24/7 Phish Fry: How Nestle built an AI-Powered phishing detector with Azure.
Language

24/7 Phish Fry: How Nestle built an AI-Powered phishing detector with Azure.

Foiling attackers who try to lure email users into clicking on a malicious link is a cat-and-mouse game, as phishing tactics evolve to evade detection. But machine learning models designed to recognize phishing attempts can evolve, too, through automatic retraining and checks to maintain accuracy.

May 12, 20212 min read
Robocoders: How SourceAI uses GPT-3 to write code in 40 languages.
Language

Robocoders: How SourceAI uses GPT-3 to write code in 40 languages.

Language models are starting to take on programming work. SourceAI uses GPT-3 to translate plain-English requests into computer code in 40 programming languages. The French startup is one of several companies that use AI to ease coding.

May 5, 20211 min read
Robocoders: How SourceAI uses GPT-3 to write code in 40 languages.
Language

Robocoders: How SourceAI uses GPT-3 to write code in 40 languages.

Language models are starting to take on programming work. SourceAI uses GPT-3 to translate plain-English requests into computer code in 40 programming languages. The French startup is one of several companies that use AI to ease coding.

May 5, 20211 min read
Up for Debate: IBM's NLP-powered debate bot mines LexisNexis.
Language

Up for Debate: IBM's NLP-powered debate bot mines LexisNexis.

IBM’s Watson question-answering system stunned the world in 2011 when it bested human champions of the TV trivia game show Jeopardy! Although the Watson brand has fallen on hard times, the company’s language-processing prowess continues to develop.

April 28, 20212 min read
Up for Debate: IBM's NLP-powered debate bot mines LexisNexis.
Language

Up for Debate: IBM's NLP-powered debate bot mines LexisNexis.

IBM’s Watson question-answering system stunned the world in 2011 when it bested human champions of the TV trivia game show Jeopardy! Although the Watson brand has fallen on hard times, the company’s language-processing prowess continues to develop.

April 28, 20212 min read
Haters Gonna [Mute]: Gamers can mute offensive language with AI.
Language

Haters Gonna [Mute]: Gamers can mute offensive language with AI.

A new tool aims to let video gamers control how much vitriol they receive from fellow players. Intel announced a voice recognition tool called Bleep that the company claims can moderate voice chat automatically, allowing users to silence offensive language.

April 21, 20212 min read
Haters Gonna [Mute]: Gamers can mute offensive language with AI.
Language

Haters Gonna [Mute]: Gamers can mute offensive language with AI.

A new tool aims to let video gamers control how much vitriol they receive from fellow players. Intel announced a voice recognition tool called Bleep that the company claims can moderate voice chat automatically, allowing users to silence offensive language.

April 21, 20212 min read
Large Language Models for Chinese: A brief overview of the WuDao family
Language

Large Language Models for Chinese: A brief overview of the WuDao family

Researchers unveiled competition for the reigning large language model GPT-3. Four models collectively called Wu Dao were described by Beijing Academy of Artificial Intelligence, a research collective funded by the Chinese government, according to Synced Review.

April 14, 20212 min read
Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.
Language

Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.

Key machine learning datasets are riddled with mistakes. Several benchmark datasets are shot through with incorrect labels. On average, 3.4 percent of examples in 10 commonly used datasets are mislabeled and the detrimental impact of such errors rises with model size.

April 14, 20212 min read
Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.
Language

Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.

Key machine learning datasets are riddled with mistakes. Several benchmark datasets are shot through with incorrect labels. On average, 3.4 percent of examples in 10 commonly used datasets are mislabeled and the detrimental impact of such errors rises with model size.

April 14, 20212 min read
Large Language Models for Chinese: A brief overview of the WuDao family
Language

Large Language Models for Chinese: A brief overview of the WuDao family

Researchers unveiled competition for the reigning large language model GPT-3. Four models collectively called Wu Dao were described by Beijing Academy of Artificial Intelligence, a research collective funded by the Chinese government, according to Synced Review.

April 14, 20212 min read
Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.
Language

Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.

It’s well established that pretraining a model on a large dataset improves performance on fine-tuned tasks. In sufficient quantity and paired with a big model, even data scraped from the internet at random can contribute to the performance boost.

April 7, 20212 min read
Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.
Language

Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.

It’s well established that pretraining a model on a large dataset improves performance on fine-tuned tasks. In sufficient quantity and paired with a big model, even data scraped from the internet at random can contribute to the performance boost.

April 7, 20212 min read
Star Trek, The Videobot Generation: William Shatner creates his own deepfake.
Language

Star Trek, The Videobot Generation: William Shatner creates his own deepfake.

A digital doppelgänger of Star Trek’s original star will let fans chat with him — possibly well beyond his lifetime. AI startup StoryFile built a lifelike videobot of actor William Shatner, best known for playing Captain James T. Kirk on Star Trek.

March 31, 20211 min read
Pictures From Words and Gestures: AI model generates captions as users mouse over images.
Language

Pictures From Words and Gestures: AI model generates captions as users mouse over images.

A new system combines verbal descriptions and crude lines to visualize complex scenes. Google researchers led by Jing Yu Koh proposed Tag-Retrieve-Compose-Synthesize (TReCS), a system that generates photorealistic images by describing what they want to see while mousing around on a blank screen.

March 31, 20212 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox