Better Agentic Prompts Automatically: Authors devised GEPA, an algorithm for better prompts to improve agentic systems’ performance
Large Language Models (LLMs)

Better Agentic Prompts Automatically: Authors devised GEPA, an algorithm for better prompts to improve agentic systems’ performance

Honing an agent’s prompt can yield better results than fine-tuning the underlying large language model via reinforcement learning.

October 22, 20253 min read
MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data
Large Language Models (LLMs)

MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data

The ability to easily connect large language models to tools and data sources has made Model Context Protocol popular among developers, but it also opens security holes, research shows.

October 22, 20252 min read
Reasoning Without “Thinking”: All about Ant Group’s Ling-1T, an open, non-reasoning model that outperforms closed competitors
Large Language Models (LLMs)

Reasoning Without “Thinking”: All about Ant Group’s Ling-1T, an open, non-reasoning model that outperforms closed competitors

Reasoning models typically learn to undertake a separate process of “thinking” through their output of before they produce final response. Ant Group built a top non-reasoning model that can take similar steps as part of its immediate response.

October 22, 20253 min read
Fine-Tuning Simplified: Thinking Machines’ new Tinker API makes it easier to fine-tune models on many GPUs
Large Language Models (LLMs)

Fine-Tuning Simplified: Thinking Machines’ new Tinker API makes it easier to fine-tune models on many GPUs

The first offering from Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, aims to simplify — and democratize — the process of fine-tuning AI models.

October 15, 20252 min read
DeepSeek Cuts Inference Costs: DeepSeek-V3.2-Exp streamlines processing using a "lightning indexer," boosting efficiency
Large Language Models (LLMs)

DeepSeek Cuts Inference Costs: DeepSeek-V3.2-Exp streamlines processing using a "lightning indexer," boosting efficiency

DeepSeek’s latest large language model can cut inference costs by more than half and processes long contexts dramatically faster relative to its predecessor.

October 15, 20253 min read
LoRA Adapters On Tap: Text-to-LoRA generates task-specific LoRA adapters directly from natural language descriptions
Large Language Models (LLMs)

LoRA Adapters On Tap: Text-to-LoRA generates task-specific LoRA adapters directly from natural language descriptions

The approach known as LoRA streamlines fine-tuning by training a small adapter that modifies a pretrained model’s weights at inference. Researchers built a model that generates such adapters directly.

October 8, 20252 min read
Qwen3 Goes Big (and Smaller): Alibaba expands Qwen3 family with a 1 trillion-parameter Max model, open-weights Qwen3-VL, and the Qwen3-Omni voice model
Large Language Models (LLMs)

Qwen3 Goes Big (and Smaller): Alibaba expands Qwen3 family with a 1 trillion-parameter Max model, open-weights Qwen3-VL, and the Qwen3-Omni voice model

Alibaba rounded out the Qwen3 family with its biggest large language model to date as well as smaller models that process text, images, video, and/or audio.

October 8, 20253 min read
OpenAI, Meta Diversify AI Product Lines: OpenAI and Meta launch social video apps while ChatGPT adds Pulse and Instant Checkout
Large Language Models (LLMs)

OpenAI, Meta Diversify AI Product Lines: OpenAI and Meta launch social video apps while ChatGPT adds Pulse and Instant Checkout

OpenAI and Meta, which have been content to offer standalone chatbots or tuck them into existing products, introduced dueling social video networks and other initiatives designed to boost revenue and engagement.

October 8, 20253 min read
Claude Levels Up: Anthropic launches Claude Sonnet 4.5 and the Claude Agent SDK, and overhauls Claude Code for developers
Large Language Models (LLMs)

Claude Levels Up: Anthropic launches Claude Sonnet 4.5 and the Claude Agent SDK, and overhauls Claude Code for developers

Anthropic updated its mid-size Claude Sonnet model, making it the first member of the Claude family to reach version 4.5. It also enhanced the Claude Code agentic coding tool with long-desired features.

October 8, 20253 min read
Faster Reinforcement Learning: New technique auto-selects training examples to speed up fine-tuning
Large Language Models (LLMs)

Faster Reinforcement Learning: New technique auto-selects training examples to speed up fine-tuning

Fine-tuning large language models via reinforcement learning is computationally expensive, but researchers found a way to streamline the process.

September 24, 20252 min read
What ChatGPT Users Want: ChatGPT users now more likely to be young, female, and seeking info, study shows
Large Language Models (LLMs)

What ChatGPT Users Want: ChatGPT users now more likely to be young, female, and seeking info, study shows

What do ChatGPT’s 700 million weekly active users do with it? OpenAI teamed up with a Harvard economist to find out.

September 24, 20253 min read
Agents of Commerce: Google’s AP2 gives developers new tools to build agentic payments
Large Language Models (LLMs)

Agents of Commerce: Google’s AP2 gives developers new tools to build agentic payments

Google launched an open protocol for agentic payments that enables agents based on any large language model to purchase items over the internet.

September 24, 20252 min read
Qwen3-Next Accelerates: Alibaba’s new model uses hybrid attention layers and a sparse MoE architecture for speed and performance
Large Language Models (LLMs)

Qwen3-Next Accelerates: Alibaba’s new model uses hybrid attention layers and a sparse MoE architecture for speed and performance

Alibaba updated its popular Qwen3 open-weights models with a number of fresh, speed-boosting tweaks.

September 17, 20253 min read
10 Million Tokens of Input Context: ATLAS, a transformer-like architecture, can process a context window as large as ten million tokens
Large Language Models (LLMs)

10 Million Tokens of Input Context: ATLAS, a transformer-like architecture, can process a context window as large as ten million tokens

An alternative to attention enables large language models to track relationships among words across extraordinarily wide spans of text.

September 10, 20253 min read
Gemini’s Environmental Impact Measured: Google study directly measures electricity, water use, and greenhouse emissions of its models
Large Language Models (LLMs)

Gemini’s Environmental Impact Measured: Google study directly measures electricity, water use, and greenhouse emissions of its models

Google determined that its large language models have a smaller environmental footprint than previous estimates had led it to expect.

September 3, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox