Learning the Language of Geometry: AlphaGeometry, a system that nears expert proficiency in proving complex geometry theorems
Machine Learning Research

Learning the Language of Geometry: AlphaGeometry, a system that nears expert proficiency in proving complex geometry theorems

Machine learning algorithms often struggle with geometry. A language model learned to prove relatively difficult theorems. 

January 24, 20242 min read
Learning the Language of Geometry: AlphaGeometry, a system that nears expert proficiency in proving complex geometry theorems
Machine Learning Research

Learning the Language of Geometry: AlphaGeometry, a system that nears expert proficiency in proving complex geometry theorems

Machine learning algorithms often struggle with geometry. A language model learned to prove relatively difficult theorems. 

January 24, 20242 min read
Sing a Tune, Generate an Accompaniment: SingSong, a tool that generates instrumental music for unaccompanied input vocals
Machine Learning Research

Sing a Tune, Generate an Accompaniment: SingSong, a tool that generates instrumental music for unaccompanied input vocals

A neural network makes music for unaccompanied vocal tracks. Chris Donahue, Antoine Caillon, Adam Roberts, and colleagues at Google proposed SingSong, a system that generates musical accompaniments for sung melodies. You can listen to its output here.

January 17, 20242 min read
Sing a Tune, Generate an Accompaniment: SingSong, a tool that generates instrumental music for unaccompanied input vocals
Machine Learning Research

Sing a Tune, Generate an Accompaniment: SingSong, a tool that generates instrumental music for unaccompanied input vocals

A neural network makes music for unaccompanied vocal tracks. Chris Donahue, Antoine Caillon, Adam Roberts, and colleagues at Google proposed SingSong, a system that generates musical accompaniments for sung melodies. You can listen to its output here.

January 17, 20242 min read
Text or Images, Input or Output: GILL, an innovative approach to multimodal model training
Machine Learning Research

Text or Images, Input or Output: GILL, an innovative approach to multimodal model training

GPT-4V introduced a large multimodal model that generates text from images and, with help from DALL-E 3, generates images from text. However, OpenAI hasn’t fully explained how it built the system. A separate group of researchers described their own method.

January 10, 20243 min read
Text or Images, Input or Output: GILL, an innovative approach to multimodal model training
Machine Learning Research

Text or Images, Input or Output: GILL, an innovative approach to multimodal model training

GPT-4V introduced a large multimodal model that generates text from images and, with help from DALL-E 3, generates images from text. However, OpenAI hasn’t fully explained how it built the system. A separate group of researchers described their own method.

January 10, 20243 min read
GPT-4 Wouldn’t Lie to Me . . . Would It?: Researchers showed how GPT-4 can deceive users without being prompted to do so explicitly
Machine Learning Research

GPT-4 Wouldn’t Lie to Me . . . Would It?: Researchers showed how GPT-4 can deceive users without being prompted to do so explicitly

It’s well known that large language models can make assertions that are blatantly false. But can they concoct outright lies? In a proof-of-concept demonstration, Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn at Apollo Research...

January 3, 20243 min read
GPT-4 Wouldn’t Lie to Me . . . Would It?: Researchers showed how GPT-4 can deceive users without being prompted to do so explicitly
Machine Learning Research

GPT-4 Wouldn’t Lie to Me . . . Would It?: Researchers showed how GPT-4 can deceive users without being prompted to do so explicitly

It’s well known that large language models can make assertions that are blatantly false. But can they concoct outright lies? In a proof-of-concept demonstration, Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn at Apollo Research...

January 3, 20243 min read
Generative AI Everywhere: How Large Language Models, chatbots, and other generative AI took off in 2023
Machine Learning Research

Generative AI Everywhere: How Large Language Models, chatbots, and other generative AI took off in 2023

This year, AI became virtually synonymous with generative AI. Launched in November 2022, OpenAI’s ChatGPT ushered in a banner year for AI-driven generation of text, images, and an ever widening range of data types. 

December 20, 20232 min read
Generative AI Everywhere: How Large Language Models, chatbots, and other generative AI took off in 2023
Machine Learning Research

Generative AI Everywhere: How Large Language Models, chatbots, and other generative AI took off in 2023

This year, AI became virtually synonymous with generative AI. Launched in November 2022, OpenAI’s ChatGPT ushered in a banner year for AI-driven generation of text, images, and an ever widening range of data types. 

December 20, 20232 min read
The Big Picture and the Details: I-JEPA, or how vision models understand the relationship between parts and the whole
Machine Learning Research

The Big Picture and the Details: I-JEPA, or how vision models understand the relationship between parts and the whole

A novel twist on self-supervised learning aims to improve on earlier methods by helping vision models learn how parts of an image relate to the whole.

December 13, 20232 min read
The Big Picture and the Details: I-JEPA, or how vision models understand the relationship between parts and the whole
Machine Learning Research

The Big Picture and the Details: I-JEPA, or how vision models understand the relationship between parts and the whole

A novel twist on self-supervised learning aims to improve on earlier methods by helping vision models learn how parts of an image relate to the whole.

December 13, 20232 min read
Multitask Vision Transformer
Machine Learning Research

Multitask Vision Transformer

The original DINO showed that a vision transformer pretrained on unlabeled images could learn representations that were sufficient for classifying and segmenting images. In an update of that work, the model learned representations useful in a wider variety of tasks.

December 6, 20233 min read
Multitask Vision Transformer
Machine Learning Research

Multitask Vision Transformer

The original DINO showed that a vision transformer pretrained on unlabeled images could learn representations that were sufficient for classifying and segmenting images. In an update of that work, the model learned representations useful in a wider variety of tasks.

December 6, 20233 min read
Robot, Find My Keys: A machine learning model for robots to predict the location of objects in households
Machine Learning Research

Robot, Find My Keys: A machine learning model for robots to predict the location of objects in households

Researchers proposed a way for robots to find objects in households where things get moved around. Andrey Kurenkov and colleagues at Stanford University introduced Node Edge Predictor, a model that learned to predict where objects were located in houses.

December 6, 20233 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox