From Prediction to Action by Tanmay Gupta: Tanmay Gupta of the Allen Institute on building AI for long-horizon tasks
Machine Learning Research

From Prediction to Action by Tanmay Gupta: Tanmay Gupta of the Allen Institute on building AI for long-horizon tasks

AI research in 2026 should confront a simple but transformative realization: Models that predict are not the same as systems that act. The latter is what we actually need.

January 2, 20263 min read
AI for Scientific Discovery by Adji Bousso Dieng: Adji Bousso Dieng, Princeton University Assistant Professor and AI Researcher, on optimizing models for the long tail
Machine Learning Research

AI for Scientific Discovery by Adji Bousso Dieng: Adji Bousso Dieng, Princeton University Assistant Professor and AI Researcher, on optimizing models for the long tail

In 2026, I hope AI will transition from being a tool for efficiency to a catalyst for scientific discovery.

January 2, 20262 min read
Open Source Wins by David Cox: David Cox, VP for AI Models at IBM Research, on the need for open development in AI
Machine Learning Research

Open Source Wins by David Cox: David Cox, VP for AI Models at IBM Research, on the need for open development in AI

My hope is that open AI continues to flourish and ultimately wins.

January 2, 20262 min read
Agents Write Code Faster, Cheaper: Software developers used more versatile AI-powered tools to write code
Machine Learning Research

Agents Write Code Faster, Cheaper: Software developers used more versatile AI-powered tools to write code

Coding apps moved beyond autofill-style code completion to agentic systems that manage a wide range of software development tasks.

December 26, 20253 min read
Thinking Models Solve Bigger Problems: Reasoning models, beginning with OpenAI’s o1 and DeepSeek’s R1, transformed the industry
Machine Learning Research

Thinking Models Solve Bigger Problems: Reasoning models, beginning with OpenAI’s o1 and DeepSeek’s R1, transformed the industry

Think step by step. Explain your reasoning. Work backwards from the answer. As 2025 began, models executed these reasoning strategies only when prompted. Now most new large language models do it as a matter of course, improving performance across a wide range of tasks.

December 26, 20253 min read
Adapting LLMs to Any Sort of Data: SEMI (Sample-Efficient Modality Integration) tackles new domains with few-shot examples
Machine Learning Research

Adapting LLMs to Any Sort of Data: SEMI (Sample-Efficient Modality Integration) tackles new domains with few-shot examples

Enabling a pretrained large language model to process a data type other than text (say, images), possibly in a specialized domain (say, radiology), typically requires thousands to millions of examples that pair the other data (perhaps x-rays) with text.

December 17, 20253 min read
OpenAI’s Answer to Gemini 3: GPT-5.2 arrives, touting variable reasoning and coding performance
Machine Learning Research

OpenAI’s Answer to Gemini 3: GPT-5.2 arrives, touting variable reasoning and coding performance

OpenAI launched GPT-5.2 only weeks after its CEO Sam Altman reportedly issued a “code red” alarm in response to Google's Gemini 3.

December 17, 20253 min read
Coherent, Interactive Worlds: Runway’s GWM-1 models generate videos with consistent physics for robots and entertainment
Machine Learning Research

Coherent, Interactive Worlds: Runway’s GWM-1 models generate videos with consistent physics for robots and entertainment

Runway’s GWM-1 family of video-generation models respond to user input in real time while producing scenes that remain consistent regardless of the camera’s position.

December 17, 20253 min read
Small Models Solve Hard Puzzles: Tiny Recursive Model beats larger competitors at games like Sudoku and Maze
Machine Learning Research

Small Models Solve Hard Puzzles: Tiny Recursive Model beats larger competitors at games like Sudoku and Maze

Large language models often fail at puzzles like Sudoku, for which a solution includes multiple elements and a single mistake invalidates all of them. Researchers showed that a tiny network, by repeatedly refining its solution, can solve this sort of puzzle well.

December 10, 20253 min read
Amazon Steps Forward: Nova 2 family boosts cost-effective performance, adds new agentic features
Machine Learning Research

Amazon Steps Forward: Nova 2 family boosts cost-effective performance, adds new agentic features

Amazon raised the competitive profile of its foundation models and added services for custom model training and an agent platform for browser automation.

December 10, 20253 min read
Claude Does More With Fewer Tokens: Claude Opus 4.5 retakes the coding crown at one-third the price of its predecessor
Machine Learning Research

Claude Does More With Fewer Tokens: Claude Opus 4.5 retakes the coding crown at one-third the price of its predecessor

Claude Opus 4.5, the latest version of Anthropic’s flagship model, extends the earlier version’s strengths in coding, computer use, and agentic workflows while generating fewer tokens.

December 10, 20253 min read
Coordinating Robot Teams: Google DeepMind’s RoboBallet project blends GNNs with RL to drive 8-armed robots
Machine Learning Research

Coordinating Robot Teams: Google DeepMind’s RoboBallet project blends GNNs with RL to drive 8-armed robots

In factories, where teams of robotic arms work in tight spaces, their motions are programmed by hand to keep them from interfering with one another. Researchers automated this programming using graph neural networks trained via reinforcement learning.

December 3, 20253 min read
Baidu’s Multimodal Bids: Giant Ernie 5 natively generates multiple media; Ernie-4.5-VL-28B-A3B-Thinking tops Vision-Language metrics
Machine Learning Research

Baidu’s Multimodal Bids: Giant Ernie 5 natively generates multiple media; Ernie-4.5-VL-28B-A3B-Thinking tops Vision-Language metrics

Baidu debuted two models: a lightweight, open-weights, vision-language model and a giant, proprietary, multimodal model built to take on U.S. competitors.

December 3, 20253 min read
Generated, Editable Virtual Spaces: World Labs makes Marble world model public, adds Chisel editing tool
Machine Learning Research

Generated, Editable Virtual Spaces: World Labs makes Marble world model public, adds Chisel editing tool

Models that generate 3D spaces typically generate them as users move through them without generating a persistent world to be explored later. A new model produces 3D worlds that can be exported and modified.

December 3, 20252 min read
Open 3D Generation Pipeline: Meta’s SAM 3 image segmentation models can analyze and create bodies and other objects
Machine Learning Research

Open 3D Generation Pipeline: Meta’s SAM 3 image segmentation models can analyze and create bodies and other objects

Meta’s Segment Anything Model (SAM) image-segmentation model has evolved into an open-weights suite for generating 3D objects. SAM 3 segments images, SAM 3D turns the segments into 3D objects, and SAM 3D Body produces 3D objects of any people among the segments. You can experiment with all three.

December 3, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox