Diffusion Models

44 Posts

Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud
Diffusion Models

Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud

The third generation of Apple Foundation Models — fruit of Apple’s collaboration with Google — introduces a variation on the mixture-of-experts architecture that runs on local devices. 

June 26, 20263 min read
Planning Generated Images In Stages: Meta improves image models by plotting and revising generations step-by-step
Diffusion Models

Planning Generated Images In Stages: Meta improves image models by plotting and revising generations step-by-step

Text-to-image generators that use diffusion or flow-matching typically compose a whole image at once (although they refine the whole image in steps).

May 29, 20264 min read
Gemini’s Music Generator: Google debuted Lyria 3, an app that turns text or images into 30-second songs
Diffusion Models

Gemini’s Music Generator: Google debuted Lyria 3, an app that turns text or images into 30-second songs

Google added a music generator to Gemini and YouTube, putting a model that produces synthetic songs in front of hundreds of millions of users.

April 3, 20263 min read
OpenAI Exits Video Generation: OpenAI will shut down Sora, its once state-of-the-art video model
Diffusion Models

OpenAI Exits Video Generation: OpenAI will shut down Sora, its once state-of-the-art video model

OpenAI plans to shut down its video generator Sora in a sudden retreat from the video market.

April 3, 20263 min read
Lightning-Fast Diffusion Learning: Inside Feature Auto-Encoder, a diffusion image generator that shrinks embeddings for more speed
Diffusion Models

Lightning-Fast Diffusion Learning: Inside Feature Auto-Encoder, a diffusion image generator that shrinks embeddings for more speed

Research shows that diffusion image generators can train somewhat faster if they learn to reconstruct embeddings from a pretrained encoder that’s built for vision tasks like classification, segmentation, and retrieval — not image generation.

March 16, 20263 min read
Refining Words in Pictures: Z.ai’s GLM-Image blends transformer and diffusion architectures for better text in images
Diffusion Models

Refining Words in Pictures: Z.ai’s GLM-Image blends transformer and diffusion architectures for better text in images

Image generators often mangle text. An open-weights model outperforms open and proprietary competitors in text rendering.

January 30, 20263 min read
Detailed Text- or Image-to-3D, Pronto: FlashWorld generates 3D objects, scenes, and surfaces with photorealistic fidelity
Diffusion Models

Detailed Text- or Image-to-3D, Pronto: FlashWorld generates 3D objects, scenes, and surfaces with photorealistic fidelity

Current methods that produce 3D scenes from text or images are slow and produce inconsistent results. Researchers introduced a technique that generates detailed, coherent 3D scenes seconds.

January 23, 20263 min read
Training Cars to Reason: Nvidia’s Alpamayo-R1 is a robotics-style reasoning model for autonomous vehicles
Diffusion Models

Training Cars to Reason: Nvidia’s Alpamayo-R1 is a robotics-style reasoning model for autonomous vehicles

Chain-of-thought reasoning can help autonomous vehicles decide what to do next.

January 23, 20262 min read
Better Images Through Reasoning: HunyuanImage-3.0 uses reinforcement learning and thinking tokens to better understand prompts
Diffusion Models

Better Images Through Reasoning: HunyuanImage-3.0 uses reinforcement learning and thinking tokens to better understand prompts

A new image generator reasons over prompts to produce outstanding pictures.

November 12, 20253 min read
Mixture of Video Experts: Alibaba’s Wan 2.2 video models adopt a new architecture to sort noisy from less-noisy inputs
Diffusion Models

Mixture of Video Experts: Alibaba’s Wan 2.2 video models adopt a new architecture to sort noisy from less-noisy inputs

The mixture-of-experts approach that has boosted the performance of large language models may do the same for video generation.

August 20, 20253 min read
Faster Learning for Diffusion Models: Pretrained embeddings accelerate diffusion transformers’ learning
Diffusion Models

Faster Learning for Diffusion Models: Pretrained embeddings accelerate diffusion transformers’ learning

Diffusion transformers learn faster when they can look at embeddings generated by a pretrained model like DINOv2.

March 26, 20252 min read
Better Images in Fewer Steps: Researchers introduce shortcut models to speed up diffusion
Diffusion Models

Better Images in Fewer Steps: Researchers introduce shortcut models to speed up diffusion

Diffusion models usually take many noise-removal steps to produce an image, which takes time at inference. There are ways to reduce the number of steps, but the resulting systems are less effective. Researchers devised a streamlined approach that doesn’t sacrifice output quality.

March 26, 20252 min read
Designer Materials: MatterGen, a diffusion model that designs new materials with specified properties
Diffusion Models

Designer Materials: MatterGen, a diffusion model that designs new materials with specified properties

Materials that have specific properties are essential to progress in critical technologies like solar cells and batteries. A machine learning model designs new materials to order.

March 19, 20252 min read
Designer Materials: MatterGen, a diffusion model that designs new materials with specified properties
Diffusion Models

Designer Materials: MatterGen, a diffusion model that designs new materials with specified properties

Materials that have specific properties are essential to progress in critical technologies like solar cells and batteries. A machine learning model designs new materials to order.

March 19, 20252 min read
Text Generation by Diffusion: Mercury Coder uses diffusion to generate text
Diffusion Models

Text Generation by Diffusion: Mercury Coder uses diffusion to generate text

Typical large language models are autoregressive, predicting the next token, one at a time, from left to right. A new model hones all text tokens at once.

March 5, 20252 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox