Text Generation by Diffusion: Mercury Coder uses diffusion to generate text
Diffusion Models

Text Generation by Diffusion: Mercury Coder uses diffusion to generate text

Typical large language models are autoregressive, predicting the next token, one at a time, from left to right. A new model hones all text tokens at once.

March 5, 20252 min read
David Ding: Generated video with music, sound effects, and dialogue
Diffusion Models

David Ding: Generated video with music, sound effects, and dialogue

Last year, we saw an explosion of models that generate either video or audio outputs in high quality. In the coming year, I look forward to models that produce video clips complete with audio soundtracks including speech, music, and sound effects.

January 1, 20253 min read
David Ding: Generated video with music, sound effects, and dialogue
Diffusion Models

David Ding: Generated video with music, sound effects, and dialogue

Last year, we saw an explosion of models that generate either video or audio outputs in high quality. In the coming year, I look forward to models that produce video clips complete with audio soundtracks including speech, music, and sound effects.

January 1, 20253 min read
Open Video Gen Closes the Gap: Tencent releases HunyuanVideo, an open source model rivaling commercial video generators
Diffusion Models

Open Video Gen Closes the Gap: Tencent releases HunyuanVideo, an open source model rivaling commercial video generators

The gap is narrowing between closed and open models for video generation.

December 18, 20242 min read
Open Video Gen Closes the Gap: Tencent releases HunyuanVideo, an open source model rivaling commercial video generators
Diffusion Models

Open Video Gen Closes the Gap: Tencent releases HunyuanVideo, an open source model rivaling commercial video generators

The gap is narrowing between closed and open models for video generation.

December 18, 20242 min read
Game Worlds on Tap: Genie 2 brings interactive 3D worlds to life
Diffusion Models

Game Worlds on Tap: Genie 2 brings interactive 3D worlds to life

A new model improves on recent progress in generating interactive virtual worlds from still images.

December 11, 20242 min read
Game Worlds on Tap: Genie 2 brings interactive 3D worlds to life
Diffusion Models

Game Worlds on Tap: Genie 2 brings interactive 3D worlds to life

A new model improves on recent progress in generating interactive virtual worlds from still images.

December 11, 20242 min read
Competitive Performance, Competitive Prices: Amazon introduces Nova models for text, image, and video
Diffusion Models

Competitive Performance, Competitive Prices: Amazon introduces Nova models for text, image, and video

Amazon introduced a range of models that confront competitors head-on.

December 11, 20244 min read
Competitive Performance, Competitive Prices: Amazon introduces Nova models for text, image, and video
Diffusion Models

Competitive Performance, Competitive Prices: Amazon introduces Nova models for text, image, and video

Amazon introduced a range of models that confront competitors head-on.

December 11, 20244 min read
Faster, Cheaper Video Generation: Pyramidal Flow Matching, a cost-cutting method for training video generators
Diffusion Models

Faster, Cheaper Video Generation: Pyramidal Flow Matching, a cost-cutting method for training video generators

Researchers devised a way to cut the cost of training video generators. They used it to build a competitive open source text-to-video model and promised to release the training code.

October 23, 20243 min read
Faster, Cheaper Video Generation: Pyramidal Flow Matching, a cost-cutting method for training video generators
Diffusion Models

Faster, Cheaper Video Generation: Pyramidal Flow Matching, a cost-cutting method for training video generators

Researchers devised a way to cut the cost of training video generators. They used it to build a competitive open source text-to-video model and promised to release the training code.

October 23, 20243 min read
For Faster Diffusion, Think a GAN: Adversarial Diffusion Distillation, a method to accelerate diffusion models
Diffusion Models

For Faster Diffusion, Think a GAN: Adversarial Diffusion Distillation, a method to accelerate diffusion models

Generative adversarial networks (GANs) produce images quickly, but they’re of relatively low quality. Diffusion image generators typically take more time, but they produce higher-quality output. Researchers aimed to achieve the best of both worlds.

June 19, 20242 min read
For Faster Diffusion, Think a GAN: Adversarial Diffusion Distillation, a method to accelerate diffusion models
Diffusion Models

For Faster Diffusion, Think a GAN: Adversarial Diffusion Distillation, a method to accelerate diffusion models

Generative adversarial networks (GANs) produce images quickly, but they’re of relatively low quality. Diffusion image generators typically take more time, but they produce higher-quality output. Researchers aimed to achieve the best of both worlds.

June 19, 20242 min read
Generative AI Calling: Google brings advanced computer vision and audio tech to Pixel 8 and 8 Pro phones.
Diffusion Models

Generative AI Calling: Google brings advanced computer vision and audio tech to Pixel 8 and 8 Pro phones.

Google’s new mobile phones put advanced computer vision and audio research into consumers’ hands. The Alphabet division introduced its flagship Pixel 8 and Pixel 8 Pro smartphones at its annual hardware-launch event. Both units feature AI-powered tools for editing photos and videos.

October 18, 20232 min read
Generative AI Calling: Google brings advanced computer vision and audio tech to Pixel 8 and 8 Pro phones.
Diffusion Models

Generative AI Calling: Google brings advanced computer vision and audio tech to Pixel 8 and 8 Pro phones.

Google’s new mobile phones put advanced computer vision and audio research into consumers’ hands. The Alphabet division introduced its flagship Pixel 8 and Pixel 8 Pro smartphones at its annual hardware-launch event. Both units feature AI-powered tools for editing photos and videos.

October 18, 20232 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox