Datasets

106 Posts

Benchmark Tests Are Meaningless: The problem with training data contamination in machine learning
Datasets

Benchmark Tests Are Meaningless: The problem with training data contamination in machine learning

The universe of web pages includes correct answers to common questions that are used to test large language models. How can we evaluate new models if they’ve studied the answers before we give them the test?

October 30, 20242 min read
Benchmark Tests Are Meaningless: The problem with training data contamination in machine learning
Datasets

Benchmark Tests Are Meaningless: The problem with training data contamination in machine learning

The universe of web pages includes correct answers to common questions that are used to test large language models. How can we evaluate new models if they’ve studied the answers before we give them the test?

October 30, 20242 min read
Balancing Web Data Distributions: Automated method organizes large datasets for better model performance
Datasets

Balancing Web Data Distributions: Automated method organizes large datasets for better model performance

Datasets that were scraped from the web tend to be unbalanced, meaning examples of some classes (say, cats) are plentiful while examples of others (say, caterpillars) are scarce.

September 11, 20243 min read
Balancing Web Data Distributions: Automated method organizes large datasets for better model performance
Datasets

Balancing Web Data Distributions: Automated method organizes large datasets for better model performance

Datasets that were scraped from the web tend to be unbalanced, meaning examples of some classes (say, cats) are plentiful while examples of others (say, caterpillars) are scarce.

September 11, 20243 min read
Data Disappears: Creative workers don't want AI developers to train models on their work
Datasets

Data Disappears: Creative workers don't want AI developers to train models on their work

The latest advances in AI are built on freely available training data. What will happen if it becomes off-limits? Creative workers don’t want AI developers to train models on their works without permission or compensation, or at all. Data is vanishing as they scramble to lock it down. 

October 25, 20232 min read
Data Disappears: Creative workers don't want AI developers to train models on their work
Datasets

Data Disappears: Creative workers don't want AI developers to train models on their work

The latest advances in AI are built on freely available training data. What will happen if it becomes off-limits? Creative workers don’t want AI developers to train models on their works without permission or compensation, or at all. Data is vanishing as they scramble to lock it down. 

October 25, 20232 min read
More Scraped Data, Greater Bias: Research shows that training on larger datasets can increase social bias.
Datasets

More Scraped Data, Greater Bias: Research shows that training on larger datasets can increase social bias.

How can we build large-scale language and vision models that don’t inherit social biases? Conventional wisdom suggests training on larger datasets, but research challenges this assumption.

October 4, 20232 min read
More Scraped Data, Greater Bias: Research shows that training on larger datasets can increase social bias.
Datasets

More Scraped Data, Greater Bias: Research shows that training on larger datasets can increase social bias.

How can we build large-scale language and vision models that don’t inherit social biases? Conventional wisdom suggests training on larger datasets, but research challenges this assumption.

October 4, 20232 min read
News Outlet Challenges AI Developers: The New York Times forbids the use of its work in training datasets.
Datasets

News Outlet Challenges AI Developers: The New York Times forbids the use of its work in training datasets.

The New York Times launched a multi-pronged attack on the use of its work in training datasets. The company updated its terms of service to forbid use of its web content and other data for training AI systems.

August 23, 20232 min read
News Outlet Challenges AI Developers: The New York Times forbids the use of its work in training datasets.
Datasets

News Outlet Challenges AI Developers: The New York Times forbids the use of its work in training datasets.

The New York Times launched a multi-pronged attack on the use of its work in training datasets. The company updated its terms of service to forbid use of its web content and other data for training AI systems.

August 23, 20232 min read
Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots
Datasets

Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots

Training an agent that controls a robot arm to perform a task — say, opening a door — that involves a sequence of motions (reach, grasp, turn, pull, release) can take from tens of thousands to millions of examples...

July 12, 20233 min read
Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots
Datasets

Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots

Training an agent that controls a robot arm to perform a task — say, opening a door — that involves a sequence of motions (reach, grasp, turn, pull, release) can take from tens of thousands to millions of examples...

July 12, 20233 min read
Stable Biases: Stable Diffusion may amplify biases in its training data.
Datasets

Stable Biases: Stable Diffusion may amplify biases in its training data.

Stable Diffusion may amplify biases in its training data in ways that promote deeply ingrained social stereotypes.

July 12, 20233 min read
Stable Biases: Stable Diffusion may amplify biases in its training data.
Datasets

Stable Biases: Stable Diffusion may amplify biases in its training data.

Stable Diffusion may amplify biases in its training data in ways that promote deeply ingrained social stereotypes.

July 12, 20233 min read
Finer Tuning: Surgical fine-tuning modifies layers based on data differences.
Datasets

Finer Tuning: Surgical fine-tuning modifies layers based on data differences.

Fine-tuning a neural network typically involves retraining every layer on new data. But research shows that networks may perform better when fine-tuning modifies only a subset of layers.

June 28, 20232 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox