Toward Open-Domain Chatbots: Meena Scores High on System for Grading NLP Chatbots
Benchmarks

Toward Open-Domain Chatbots: Meena Scores High on System for Grading NLP Chatbots

Progress in language models is spawning a new breed of chatbots and, unlike their narrow-domain forebears, they have the gift of gab. Recent research tests the limits of conversational AI.

May 27, 20202 min read
Toward Open-Domain Chatbots: Meena Scores High on System for Grading NLP Chatbots
Benchmarks

Toward Open-Domain Chatbots: Meena Scores High on System for Grading NLP Chatbots

Progress in language models is spawning a new breed of chatbots and, unlike their narrow-domain forebears, they have the gift of gab. Recent research tests the limits of conversational AI.

May 27, 20202 min read
What Language Models Know
Benchmarks

What Language Models Know

Watson set a high bar for language understanding in 2011, when it famously whipped human competitors in the televised trivia game show Jeopardy! IBM’s special-purpose AI required around $1 billion. Research suggests that today’s best language models can accomplish similar tasks right off the shelf.

September 11, 20192 min read
Leveling the Playing Field
Benchmarks

Leveling the Playing Field

Deep reinforcement learning has given machines apparent hegemony in vintage Atari games, but their scores have been hard to compare — with one another or with human performance — because there are no rules governing what machines can and can’t do to win. Researchers aim to change that.

September 11, 20192 min read
What Language Models Know
Benchmarks

What Language Models Know

Watson set a high bar for language understanding in 2011, when it famously whipped human competitors in the televised trivia game show Jeopardy! IBM’s special-purpose AI required around $1 billion. Research suggests that today’s best language models can accomplish similar tasks right off the shelf.

September 11, 20192 min read
Leveling the Playing Field
Benchmarks

Leveling the Playing Field

Deep reinforcement learning has given machines apparent hegemony in vintage Atari games, but their scores have been hard to compare — with one another or with human performance — because there are no rules governing what machines can and can’t do to win. Researchers aim to change that.

September 11, 20192 min read
BERT Is Back
Benchmarks

BERT Is Back

Less than a month after XLNet overtook BERT, the pole position in natural language understanding changed hands again. RoBERTa is an improved BERT pretraining recipe that beats its forbear, becoming the new state-of-the-art language model — for the moment.

August 21, 20192 min read
BERT Is Back
Benchmarks

BERT Is Back

Less than a month after XLNet overtook BERT, the pole position in natural language understanding changed hands again. RoBERTa is an improved BERT pretraining recipe that beats its forbear, becoming the new state-of-the-art language model — for the moment.

August 21, 20192 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox