Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems
Datasets

Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems

Images in the wild may not come with labels, but they often include metadata. A new training method takes advantage of this information to improve contrastive learning.

July 29, 20222 min read
Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems
Datasets

Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems

Images in the wild may not come with labels, but they often include metadata. A new training method takes advantage of this information to improve contrastive learning.

July 29, 20222 min read
Yoav Shoham: Language models that reason
Datasets

Yoav Shoham: Language models that reason

I believe that natural language processing in 2022 will re-embrace symbolic reasoning, harmonizing it with the statistical operation of modern neural networks. Let me explain what I mean by this.

December 29, 20212 min read
Yoav Shoham: Language models that reason
Datasets

Yoav Shoham: Language models that reason

I believe that natural language processing in 2022 will re-embrace symbolic reasoning, harmonizing it with the statistical operation of modern neural networks. Let me explain what I mean by this.

December 29, 20212 min read
Chip Huyen: AI that adapts to changing conditions
Datasets

Chip Huyen: AI that adapts to changing conditions

Until recently, big data processing has been dominated by batch systems like MapReduce and Spark, which allow us to periodically process a large amount of data very efficiently.

December 29, 20212 min read
Chip Huyen: AI that adapts to changing conditions
Datasets

Chip Huyen: AI that adapts to changing conditions

Until recently, big data processing has been dominated by batch systems like MapReduce and Spark, which allow us to periodically process a large amount of data very efficiently.

December 29, 20212 min read
Abeba Birhane: Clean up web datasets
Datasets

Abeba Birhane: Clean up web datasets

From language to vision models, deep neural networks are marked by improved performance, higher efficiency, and better generalizations. Yet, these systems are also marked by perpetuation of bias and injustice.

December 29, 20213 min read
Abeba Birhane: Clean up web datasets
Datasets

Abeba Birhane: Clean up web datasets

From language to vision models, deep neural networks are marked by improved performance, higher efficiency, and better generalizations. Yet, these systems are also marked by perpetuation of bias and injustice.

December 29, 20213 min read
New Models Inherit Old Flaws: Foundation models can pass along biases to their  fine-tuned progeny.
Datasets

New Models Inherit Old Flaws: Foundation models can pass along biases to their fine-tuned progeny.

Is AI becoming inbred? The fear: The best models increasingly are fine-tuned versions of a small number of so-called foundation models that were pretrained on immense quantities of data scraped from the web.

October 27, 20211 min read
New Models Inherit Old Flaws: Foundation models can pass along biases to their  fine-tuned progeny.
Datasets

New Models Inherit Old Flaws: Foundation models can pass along biases to their fine-tuned progeny.

Is AI becoming inbred? The fear: The best models increasingly are fine-tuned versions of a small number of so-called foundation models that were pretrained on immense quantities of data scraped from the web.

October 27, 20211 min read
Crawl the Web, Absorb the Bias: NLP Models Absorb Biases from Web Training Data
Datasets

Crawl the Web, Absorb the Bias: NLP Models Absorb Biases from Web Training Data

The emerging generation of trillion-parameter models needs datasets of billions of examples, but the most readily available source of examples on that scale — the web — is polluted with bias and antisocial expressions. A new study examines the issue.

October 13, 20212 min read
Crawl the Web, Absorb the Bias: NLP Models Absorb Biases from Web Training Data
Datasets

Crawl the Web, Absorb the Bias: NLP Models Absorb Biases from Web Training Data

The emerging generation of trillion-parameter models needs datasets of billions of examples, but the most readily available source of examples on that scale — the web — is polluted with bias and antisocial expressions. A new study examines the issue.

October 13, 20212 min read
Biomedical Treasure Chest: DeepMind open sources AlphaFold and protein databases.
Datasets

Biomedical Treasure Chest: DeepMind open sources AlphaFold and protein databases.

DeepMind opened access to AlphaFold, a model that finds the shapes of proteins, and to its output so far — a potential cornucopia for biomedical research. The research lab, a division of Google’s parent company Alphabet, made AlphaFold freely available.

August 4, 20212 min read
Biomedical Treasure Chest: DeepMind open sources AlphaFold and protein databases.
Datasets

Biomedical Treasure Chest: DeepMind open sources AlphaFold and protein databases.

DeepMind opened access to AlphaFold, a model that finds the shapes of proteins, and to its output so far — a potential cornucopia for biomedical research. The research lab, a division of Google’s parent company Alphabet, made AlphaFold freely available.

August 4, 20212 min read
Bias By the Book: Researchers find bias in influential NLP dataset BookCorpus.
Datasets

Bias By the Book: Researchers find bias in influential NLP dataset BookCorpus.

Researchers found serious flaws in an influential language dataset, highlighting the need for better documentation of data used in machine learning.

May 26, 20211 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox