Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.
Language

Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.

Let’s say you have a pretrained language model and a small amount of data to fine-tune it to answer yes-or-no questions. Should you fine-tune it to classify yes/no or to fill in missing words — both viable approaches that are likely to yield different results?

February 23, 20223 min read
Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.
Language

Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.

Let’s say you have a pretrained language model and a small amount of data to fine-tune it to answer yes-or-no questions. Should you fine-tune it to classify yes/no or to fill in missing words — both viable approaches that are likely to yield different results?

February 23, 20223 min read
Competitive Coder: AI code writing system can compete alongside humans.
Language

Competitive Coder: AI code writing system can compete alongside humans.

Programming is hard. Programming competitions are harder. Yet transformers proved themselves up to the task.

February 9, 20222 min read
New Supercomputer on the Block: All about Meta's AI Research Supercluster
Language

New Supercomputer on the Block: All about Meta's AI Research Supercluster

Facebook’s parent company is staking its future on a new compute cluster. Meta unveiled AI Research SuperCluster (RSC), which is designed to accelerate training of large models for applications like computer vision, natural language processing, and speech recognition.

February 9, 20222 min read
Competitive Coder: AI code writing system can compete alongside humans.
Language

Competitive Coder: AI code writing system can compete alongside humans.

Programming is hard. Programming competitions are harder. Yet transformers proved themselves up to the task.

February 9, 20222 min read
New Supercomputer on the Block: All about Meta's AI Research Supercluster
Language

New Supercomputer on the Block: All about Meta's AI Research Supercluster

Facebook’s parent company is staking its future on a new compute cluster. Meta unveiled AI Research SuperCluster (RSC), which is designed to accelerate training of large models for applications like computer vision, natural language processing, and speech recognition.

February 9, 20222 min read
A Kinder, Gentler Language Model: Inside Instruct GPT-3, OpenAI's GPT-3 successor.
Language

A Kinder, Gentler Language Model: Inside Instruct GPT-3, OpenAI's GPT-3 successor.

OpenAI unveiled a more reliable successor to its GPT-3 natural language model. InstructGPT is a version of GPT-3 fine-tuned to minimize harmful, untruthful, and biased output. It's available via an application programming interface.

February 2, 20222 min read
A Kinder, Gentler Language Model: Inside Instruct GPT-3, OpenAI's GPT-3 successor.
Language

A Kinder, Gentler Language Model: Inside Instruct GPT-3, OpenAI's GPT-3 successor.

OpenAI unveiled a more reliable successor to its GPT-3 natural language model. InstructGPT is a version of GPT-3 fine-tuned to minimize harmful, untruthful, and biased output. It's available via an application programming interface.

February 2, 20222 min read
More Learning With Less Memory: Training large language models using less memory.
Language

More Learning With Less Memory: Training large language models using less memory.

Researchers discovered a new way to reduce memory requirements when training large machine learning models. Tim Dettmers and colleagues at University of Washington released 8-bit optimizers that store gradient statistics as 8-bit values, while maintaining the same accuracy.

January 12, 20222 min read
AI Versus the Garbage Heap: How Amazon uses AI to cut waste.
Language

AI Versus the Garbage Heap: How Amazon uses AI to cut waste.

Amazon reported long-term success using machine learning to shrink its environmental footprint. The online retailer developed a system that fuses product descriptions, images, and structured data to decide how an item should be packed for shipping.

January 12, 20222 min read
More Learning With Less Memory: Training large language models using less memory.
Language

More Learning With Less Memory: Training large language models using less memory.

Researchers discovered a new way to reduce memory requirements when training large machine learning models. Tim Dettmers and colleagues at University of Washington released 8-bit optimizers that store gradient statistics as 8-bit values, while maintaining the same accuracy.

January 12, 20222 min read
AI Versus the Garbage Heap: How Amazon uses AI to cut waste.
Language

AI Versus the Garbage Heap: How Amazon uses AI to cut waste.

Amazon reported long-term success using machine learning to shrink its environmental footprint. The online retailer developed a system that fuses product descriptions, images, and structured data to decide how an item should be packed for shipping.

January 12, 20222 min read
Yoav Shoham: Language models that reason
Language

Yoav Shoham: Language models that reason

I believe that natural language processing in 2022 will re-embrace symbolic reasoning, harmonizing it with the statistical operation of modern neural networks. Let me explain what I mean by this.

December 29, 20212 min read
Yoav Shoham: Language models that reason
Language

Yoav Shoham: Language models that reason

I believe that natural language processing in 2022 will re-embrace symbolic reasoning, harmonizing it with the statistical operation of modern neural networks. Let me explain what I mean by this.

December 29, 20212 min read
Abeba Birhane: Clean up web datasets
Language

Abeba Birhane: Clean up web datasets

From language to vision models, deep neural networks are marked by improved performance, higher efficiency, and better generalizations. Yet, these systems are also marked by perpetuation of bias and injustice.

December 29, 20213 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox