Finer Tuning: Surgical fine-tuning modifies layers based on data differences.
Datasets

Finer Tuning: Surgical fine-tuning modifies layers based on data differences.

Fine-tuning a neural network typically involves retraining every layer on new data. But research shows that networks may perform better when fine-tuning modifies only a subset of layers.

June 28, 20232 min read
Training Data Free-For-All: Japan's AI data laws, explained
Datasets

Training Data Free-For-All: Japan's AI data laws, explained

Amid rising questions about the fairness and legality of using publicly available information to train AI models, Japan affirmed that machine learning engineers can use any data they find.

June 14, 20232 min read
Training Data Free-For-All: Japan's AI data laws, explained
Datasets

Training Data Free-For-All: Japan's AI data laws, explained

Amid rising questions about the fairness and legality of using publicly available information to train AI models, Japan affirmed that machine learning engineers can use any data they find.

June 14, 20232 min read
LAION Roars: The story of LAION, the dataset behind Stable Diffusion
Datasets

LAION Roars: The story of LAION, the dataset behind Stable Diffusion

The largest dataset for training text-to-image generators was assembled by volunteers for roughly $10,000. Now it’s implicated in fights over whether copyrighted works can be used for training.

June 7, 20233 min read
LAION Roars: The story of LAION, the dataset behind Stable Diffusion
Datasets

LAION Roars: The story of LAION, the dataset behind Stable Diffusion

The largest dataset for training text-to-image generators was assembled by volunteers for roughly $10,000. Now it’s implicated in fights over whether copyrighted works can be used for training.

June 7, 20233 min read
Data Does Not Want to Be Free: Reddit and Stack Overflow ask AI devs to pay for data.
Datasets

Data Does Not Want to Be Free: Reddit and Stack Overflow ask AI devs to pay for data.

Developers of language models will have to pay for access to troves of text data that they previously got for free. The discussion platform Reddit and question-and-answer site Stack Overflow announced plans to protect their data from being used to train large language models.

April 26, 20232 min read
Data Does Not Want to Be Free: Reddit and Stack Overflow ask AI devs to pay for data.
Datasets

Data Does Not Want to Be Free: Reddit and Stack Overflow ask AI devs to pay for data.

Developers of language models will have to pay for access to troves of text data that they previously got for free. The discussion platform Reddit and question-and-answer site Stack Overflow announced plans to protect their data from being used to train large language models.

April 26, 20232 min read
PCA Raises Red Flags: Principal component analysis can negatively impact science.
Datasets

PCA Raises Red Flags: Principal component analysis can negatively impact science.

Principal component analysis is a key machine learning technique for reducing the number of dimensions in a dataset, but new research shows that its output can be inconsistent and unreliable.

March 2, 20232 min read
PCA Raises Red Flags: Principal component analysis can negatively impact science.
Datasets

PCA Raises Red Flags: Principal component analysis can negatively impact science.

Principal component analysis is a key machine learning technique for reducing the number of dimensions in a dataset, but new research shows that its output can be inconsistent and unreliable.

March 2, 20232 min read
Unsupervised Data Pruning: New method removes useless machine learning data.
Datasets

Unsupervised Data Pruning: New method removes useless machine learning data.

Large datasets often contain overly similar examples that consume training cycles without contributing to learning. A new paper identifies similar training examples, even if they’re not labeled.

February 15, 20232 min read
Unsupervised Data Pruning: New method removes useless machine learning data.
Datasets

Unsupervised Data Pruning: New method removes useless machine learning data.

Large datasets often contain overly similar examples that consume training cycles without contributing to learning. A new paper identifies similar training examples, even if they’re not labeled.

February 15, 20232 min read
Language Models Defy Logic: Large NLP models struggle with logical reasoning.
Datasets

Language Models Defy Logic: Large NLP models struggle with logical reasoning.

Who would disagree that, if all people are mortal and Socrates is a person, Socrates must be mortal? GPT-3, for one. Recent work shows that bigger language models are not necessarily better when it comes to logical reasoning.

February 1, 20232 min read
Language Models Defy Logic: Large NLP models struggle with logical reasoning.
Datasets

Language Models Defy Logic: Large NLP models struggle with logical reasoning.

Who would disagree that, if all people are mortal and Socrates is a person, Socrates must be mortal? GPT-3, for one. Recent work shows that bigger language models are not necessarily better when it comes to logical reasoning.

February 1, 20232 min read
Will We Have Enough Data?
Datasets

Will We Have Enough Data?

The world’s supply of data soon may fail to meet the demands of increasingly hungry machine learning models. Researchers at Epoch AI found that a shortage of text data could cause trouble as early as this year. Vision data may fall short within a decade.

January 4, 20233 min read
Will We Have Enough Data?
Datasets

Will We Have Enough Data?

The world’s supply of data soon may fail to meet the demands of increasingly hungry machine learning models. Researchers at Epoch AI found that a shortage of text data could cause trouble as early as this year. Vision data may fall short within a decade.

January 4, 20233 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox