Bias By the Book: Researchers find bias in influential NLP dataset BookCorpus.
Datasets

Bias By the Book: Researchers find bias in influential NLP dataset BookCorpus.

Researchers found serious flaws in an influential language dataset, highlighting the need for better documentation of data used in machine learning.

May 26, 20211 min read
Boosting Biomedicine: The NIH's Bridge2AI program will fund new health datasets.
Datasets

Boosting Biomedicine: The NIH's Bridge2AI program will fund new health datasets.

The U.S. government aims to turbocharge biomedical AI research. The National Institutes of Health, which invests $41.7 billion annually in medical research, announced a program called Bridge to Artificial Intelligence (Bridge2AI) to promote machine learning in human biology and medicine.

April 28, 20211 min read
Boosting Biomedicine: The NIH's Bridge2AI program will fund new health datasets.
Datasets

Boosting Biomedicine: The NIH's Bridge2AI program will fund new health datasets.

The U.S. government aims to turbocharge biomedical AI research. The National Institutes of Health, which invests $41.7 billion annually in medical research, announced a program called Bridge to Artificial Intelligence (Bridge2AI) to promote machine learning in human biology and medicine.

April 28, 20211 min read
Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.
Datasets

Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.

Key machine learning datasets are riddled with mistakes. Several benchmark datasets are shot through with incorrect labels. On average, 3.4 percent of examples in 10 commonly used datasets are mislabeled and the detrimental impact of such errors rises with model size.

April 14, 20212 min read
Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.
Datasets

Labeling Errors Everywhere: Many deep learning datasets contain mislabeled data.

Key machine learning datasets are riddled with mistakes. Several benchmark datasets are shot through with incorrect labels. On average, 3.4 percent of examples in 10 commonly used datasets are mislabeled and the detrimental impact of such errors rises with model size.

April 14, 20212 min read
De-Facing ImageNet: Researchers blur all faces in ImageNet.
Datasets

De-Facing ImageNet: Researchers blur all faces in ImageNet.

ImageNet now comes with privacy protection.What’s new: The team that manages the machine learning community’s go-to image dataset blurred all the human faces pictured in it and tested how models trained on the modified images on a variety of image recognition tasks.

April 7, 20212 min read
De-Facing ImageNet: Researchers blur all faces in ImageNet.
Datasets

De-Facing ImageNet: Researchers blur all faces in ImageNet.

ImageNet now comes with privacy protection.What’s new: The team that manages the machine learning community’s go-to image dataset blurred all the human faces pictured in it and tested how models trained on the modified images on a variety of image recognition tasks.

April 7, 20212 min read
Spotlight on Unreproducible Results: Papers Without Code collects unreproducible AI research.
Datasets

Spotlight on Unreproducible Results: Papers Without Code collects unreproducible AI research.

A new website calls out AI research that may not lend itself to being reproduced. Papers Without Code maintains a directory of AI systems that researchers tried but failed to reproduce.

March 24, 20212 min read
Spotlight on Unreproducible Results: Papers Without Code collects unreproducible AI research.
Datasets

Spotlight on Unreproducible Results: Papers Without Code collects unreproducible AI research.

A new website calls out AI research that may not lend itself to being reproduced. Papers Without Code maintains a directory of AI systems that researchers tried but failed to reproduce.

March 24, 20212 min read
ImageNet Performance, No Panacea: ImageNet pretraining won't always improve computer vision.
Datasets

ImageNet Performance, No Panacea: ImageNet pretraining won't always improve computer vision.

It’s commonly assumed that models pretrained to achieve high performance on ImageNet will perform better on other visual tasks after fine-tuning. But is it always true? A new study reached surprising conclusions.

March 10, 20212 min read
ImageNet Performance, No Panacea: ImageNet pretraining won't always improve computer vision.
Datasets

ImageNet Performance, No Panacea: ImageNet pretraining won't always improve computer vision.

It’s commonly assumed that models pretrained to achieve high performance on ImageNet will perform better on other visual tasks after fine-tuning. But is it always true? A new study reached surprising conclusions.

March 10, 20212 min read
Cutting Corners to Recognize Faces: Research finds flaws in face recognition datasets.
Datasets

Cutting Corners to Recognize Faces: Research finds flaws in face recognition datasets.

Datasets for training face recognition models have ballooned in size — while slipping in quality and respect for privacy. In a survey of 130 datasets compiled over the last four decades, researchers traced how the need for increasing quantities of data led researchers to relax their standards.

February 24, 20212 min read
Cutting Corners to Recognize Faces: Research finds flaws in face recognition datasets.
Datasets

Cutting Corners to Recognize Faces: Research finds flaws in face recognition datasets.

Datasets for training face recognition models have ballooned in size — while slipping in quality and respect for privacy. In a survey of 130 datasets compiled over the last four decades, researchers traced how the need for increasing quantities of data led researchers to relax their standards.

February 24, 20212 min read
Facing Failure to Generalize: Why some AI models exhibit underspecification.
Datasets

Facing Failure to Generalize: Why some AI models exhibit underspecification.

The same models trained on the same data may show the same performance in the lab, and yet respond very differently to data they haven’t seen before. New work finds this inconsistency to be pervasive.

February 17, 20212 min read
Facing Failure to Generalize: Why some AI models exhibit underspecification.
Datasets

Facing Failure to Generalize: Why some AI models exhibit underspecification.

The same models trained on the same data may show the same performance in the lab, and yet respond very differently to data they haven’t seen before. New work finds this inconsistency to be pervasive.

February 17, 20212 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox