Masked Pretraining for CNNs: ConvNeXt V2, the new model family that boosts ConvNet performance
Machine Learning Research

Masked Pretraining for CNNs: ConvNeXt V2, the new model family that boosts ConvNet performance

Vision transformers have bested convolutional neural networks (CNNs) in a number of key vision tasks. Have CNNs hit their limit? New research suggests otherwise.

September 13, 20232 min read
Different Media, Similar Embeddings: ImageBind, the AI model that binds data from seven data types at once
Machine Learning Research

Different Media, Similar Embeddings: ImageBind, the AI model that binds data from seven data types at once

The ability of OpenAI’s CLIP to produce similar embeddings of a text phrase and a matching image opened up applications like classifying images according to labels that weren’t in the training set. A new model extends this capability to seven data types.

September 6, 20232 min read
Different Media, Similar Embeddings: ImageBind, the AI model that binds data from seven data types at once
Machine Learning Research

Different Media, Similar Embeddings: ImageBind, the AI model that binds data from seven data types at once

The ability of OpenAI’s CLIP to produce similar embeddings of a text phrase and a matching image opened up applications like classifying images according to labels that weren’t in the training set. A new model extends this capability to seven data types.

September 6, 20232 min read
Text-To-3D Animation: MAV3D, a method for generating 3D dynamic scenes from text descriptions
Machine Learning Research

Text-To-3D Animation: MAV3D, a method for generating 3D dynamic scenes from text descriptions

Text-to-video generation is so 2022! A new system takes in text and generates an animated 3D scene that can be viewed or rendered from any angle.

August 30, 20233 min read
Text-To-3D Animation: MAV3D, a method for generating 3D dynamic scenes from text descriptions
Machine Learning Research

Text-To-3D Animation: MAV3D, a method for generating 3D dynamic scenes from text descriptions

Text-to-video generation is so 2022! A new system takes in text and generates an animated 3D scene that can be viewed or rendered from any angle.

August 30, 20233 min read
Vision Transformers Made Manageable: FlexiViT, the vision transformer that allows users to specify the patch size
Machine Learning Research

Vision Transformers Made Manageable: FlexiViT, the vision transformer that allows users to specify the patch size

Vision transformers typically process images in patches of fixed size. Smaller patches yield higher accuracy but require more computation. A new training method lets AI engineers adjust the tradeoff.

August 23, 20232 min read
Vision Transformers Made Manageable: FlexiViT, the vision transformer that allows users to specify the patch size
Machine Learning Research

Vision Transformers Made Manageable: FlexiViT, the vision transformer that allows users to specify the patch size

Vision transformers typically process images in patches of fixed size. Smaller patches yield higher accuracy but require more computation. A new training method lets AI engineers adjust the tradeoff.

August 23, 20232 min read
LLMs Get a Life: The generative agents that mimic human behavior in a simulated town
Machine Learning Research

LLMs Get a Life: The generative agents that mimic human behavior in a simulated town

Large language models increasingly reply to prompts with a believably human response. Can they also mimic human behavior?

August 16, 20233 min read
LLMs Get a Life: The generative agents that mimic human behavior in a simulated town
Machine Learning Research

LLMs Get a Life: The generative agents that mimic human behavior in a simulated town

Large language models increasingly reply to prompts with a believably human response. Can they also mimic human behavior?

August 16, 20233 min read
Diffusion Transformed: A new class of diffusion models based on the transformer architecture
Machine Learning Research

Diffusion Transformed: A new class of diffusion models based on the transformer architecture

A tweak to diffusion models, which are responsible for most of the recent excitement about AI-generated images, enables them to produce more realistic output.

August 9, 20232 min read
Diffusion Transformed: A new class of diffusion models based on the transformer architecture
Machine Learning Research

Diffusion Transformed: A new class of diffusion models based on the transformer architecture

A tweak to diffusion models, which are responsible for most of the recent excitement about AI-generated images, enables them to produce more realistic output.

August 9, 20232 min read
Long-Range Weather Forecasts: This ML-based forecast simulator outperformed medium-range forecast systems.
Machine Learning Research

Long-Range Weather Forecasts: This ML-based forecast simulator outperformed medium-range forecast systems.

Machine learning models have predicted weather a few days ahead of time. A new approach substantially extends the time horizon. Remi Lam and colleagues at Google developed GraphCast, a weather-forecasting system based on graph neural networks (GNNs).

August 2, 20234 min read
Long-Range Weather Forecasts: This ML-based forecast simulator outperformed medium-range forecast systems.
Machine Learning Research

Long-Range Weather Forecasts: This ML-based forecast simulator outperformed medium-range forecast systems.

Machine learning models have predicted weather a few days ahead of time. A new approach substantially extends the time horizon. Remi Lam and colleagues at Google developed GraphCast, a weather-forecasting system based on graph neural networks (GNNs).

August 2, 20234 min read
Stratego Master: DeepNash, the RL system that plays Stratego like a master
Machine Learning Research

Stratego Master: DeepNash, the RL system that plays Stratego like a master

Reinforcement learning agents have mastered games like Go that provide complete information about the state of the game to players. They’ve also excelled at Texas Hold ’Em poker, which provides incomplete information, as few cards are revealed.

July 26, 20233 min read
Stratego Master: DeepNash, the RL system that plays Stratego like a master
Machine Learning Research

Stratego Master: DeepNash, the RL system that plays Stratego like a master

Reinforcement learning agents have mastered games like Go that provide complete information about the state of the game to players. They’ve also excelled at Texas Hold ’Em poker, which provides incomplete information, as few cards are revealed.

July 26, 20233 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox