Toward Steering LLM Personality: Persona Vectors allow model builders to identify and edit out sycophancy, hallucinations, and more
Machine Learning Research

Toward Steering LLM Personality: Persona Vectors allow model builders to identify and edit out sycophancy, hallucinations, and more

Large language models can develop character traits like cheerfulness or sycophancy during fine-tuning. Researchers developed a method to identify, monitor, and control such traits.

November 26, 20253 min read
Google Dominates Arena Leaderboards (For the Moment): Gemini 3 Pro and Nano Banana Pro boast best-in-class multimodal reasoning and image generation
Machine Learning Research

Google Dominates Arena Leaderboards (For the Moment): Gemini 3 Pro and Nano Banana Pro boast best-in-class multimodal reasoning and image generation

Google introduced Gemini 3 Pro and Nano Banana Pro, its flagship vision-language and image-generation models, and deployed them to billions of users worldwide.

November 26, 20254 min read
More-Efficient Agentic Search: Researchers fine-tune models to search their own parameters to boost recall
Machine Learning Research

More-Efficient Agentic Search: Researchers fine-tune models to search their own parameters to boost recall

Large language models may have learned knowledge that’s relevant to a given prompt, but they don’t always recall it consistently. Fine-tuning a model to search its parameters as though it were searching the web can help it find knowledge in its own weights.

November 19, 20253 min read
Anthropic Cyberattack Report Sparks Controversy: Security researchers question whether coding agents allow unprecedented automated attacks
Machine Learning Research

Anthropic Cyberattack Report Sparks Controversy: Security researchers question whether coding agents allow unprecedented automated attacks

Independent cybersecurity researchers pushed back on a report by Anthropic that claimed hackers had used its Claude Code agentic coding system to perpetrate an unprecedented automated cyberattack.

November 19, 20253 min read
Top Agentic Results, Open Weights: Kimi K2 Thinking outperforms proprietary models with new techniques for agentic tool use
Machine Learning Research

Top Agentic Results, Open Weights: Kimi K2 Thinking outperforms proprietary models with new techniques for agentic tool use

The latest open-weights large language model from Moonshot AI challenges top proprietary LLMs at agentic tasks by executing hundreds of tool calls sequentially and pausing to think between each.

November 19, 20254 min read
Self-Driving Cars on U.S. Freeways: Waymo deploys autonomous cars on California and Arizona expressways
Machine Learning Research

Self-Driving Cars on U.S. Freeways: Waymo deploys autonomous cars on California and Arizona expressways

Waymo became the first company to offer fully autonomous, driverless taxi service on freeways in the United States.

November 19, 20253 min read
Forecasting Multiple Time Series: Amazon’s Chronos-2 sorts out tangled variables to make better predictions
Machine Learning Research

Forecasting Multiple Time Series: Amazon’s Chronos-2 sorts out tangled variables to make better predictions

Transformers are well suited to predicting future values of time series like energy prices, wages, or weather, but often — as in those examples — multiple time series often influence one another. Researchers built a model that can forecast multiple time series simultaneously.

November 12, 20253 min read
The Year AI Went Industrial: The State of AI Report 2025 says AI’s barriers aren’t technological but social and material
Machine Learning Research

The Year AI Went Industrial: The State of AI Report 2025 says AI’s barriers aren’t technological but social and material

A year-in-review report heralds the dawn of AI’s industrial era.

November 12, 20253 min read
Better Images Through Reasoning: HunyuanImage-3.0 uses reinforcement learning and thinking tokens to better understand prompts
Machine Learning Research

Better Images Through Reasoning: HunyuanImage-3.0 uses reinforcement learning and thinking tokens to better understand prompts

A new image generator reasons over prompts to produce outstanding pictures.

November 12, 20253 min read
Masking Private Data in Training Sets: Google researchers released VaultGemma, an open-weights model redacting personal information
Machine Learning Research

Masking Private Data in Training Sets: Google researchers released VaultGemma, an open-weights model redacting personal information

Large language models often memorize details in their training data, including private information that may appear only once, like a person’s name, address, or phone number. Researchers built the first open-weights language model that’s guaranteed not to remember such facts.

November 5, 20253 min read
Open-Weights Coding Leader: MiniMax-M2’s lightweight footprint and low costs belie that its top performance
Machine Learning Research

Open-Weights Coding Leader: MiniMax-M2’s lightweight footprint and low costs belie that its top performance

An open-weights model from Shanghai-based MiniMax challenges top proprietary models on key benchmarks for coding and agentic tasks.

November 5, 20253 min read
Web Data Diminishes: What if online publishers make it harder and more expensive to train models?
Machine Learning Research

Web Data Diminishes: What if online publishers make it harder and more expensive to train models?

For decades, AI developers have treated the web as an open faucet of training data. Now publishers are shutting the tap. Will web data dry up?

October 29, 20253 min read
Better Agentic Prompts Automatically: Authors devised GEPA, an algorithm for better prompts to improve agentic systems’ performance
Machine Learning Research

Better Agentic Prompts Automatically: Authors devised GEPA, an algorithm for better prompts to improve agentic systems’ performance

Honing an agent’s prompt can yield better results than fine-tuning the underlying large language model via reinforcement learning.

October 22, 20253 min read
MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data
Machine Learning Research

MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data

The ability to easily connect large language models to tools and data sources has made Model Context Protocol popular among developers, but it also opens security holes, research shows.

October 22, 20252 min read
Reasoning Without “Thinking”: All about Ant Group’s Ling-1T, an open, non-reasoning model that outperforms closed competitors
Machine Learning Research

Reasoning Without “Thinking”: All about Ant Group’s Ling-1T, an open, non-reasoning model that outperforms closed competitors

Reasoning models typically learn to undertake a separate process of “thinking” through their output of before they produce final response. Ant Group built a top non-reasoning model that can take similar steps as part of its immediate response.

October 22, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox