GPT-5.4’s Higher Performance, Higher Price: OpenAI’s GPT-5.4 Pro and GPT-5.4 Thinking challenge Google’s Gemini 3.1 Pro Preview as best all-around AI model
Large Language Models (LLMs)

GPT-5.4’s Higher Performance, Higher Price: OpenAI’s GPT-5.4 Pro and GPT-5.4 Thinking challenge Google’s Gemini 3.1 Pro Preview as best all-around AI model

OpenAI updated its flagship models, extending the ability to use tools and setting the state of the art on a handful of benchmarks, and priced them at the top of the market. Its coding and agentic abilities have enabled Codex, OpenAI’s competitor to Anthropic’s Claude Code, to leap ahead.

March 16, 20263 min read
Agent Solves Stubborn Math Problems: Google’s Aletheia uses Gemini 3 Deep Think to find original mathematics solutions
Large Language Models (LLMs)

Agent Solves Stubborn Math Problems: Google’s Aletheia uses Gemini 3 Deep Think to find original mathematics solutions

LLMs have achieved gold-medal performance in math competitions. An agentic system showed strength in mathematical research as well.

March 6, 20263 min read
Can Local AI Stand In for the Cloud?: Stanford and Together.AI researchers chart edge models’ performance in intelligence per watt
Large Language Models (LLMs)

Can Local AI Stand In for the Cloud?: Stanford and Together.AI researchers chart edge models’ performance in intelligence per watt

Projected demand for output from large language models is spurring a massive buildout of data centers. Researchers asked whether smaller models running on local devices could meaningfully lighten that load.

February 27, 20263 min read
Investors Panic Over Agentic AI: Claude Cowork plugins trigger a SaaS stock selloff, but partnerships lead to slight rebound
Large Language Models (LLMs)

Investors Panic Over Agentic AI: Claude Cowork plugins trigger a SaaS stock selloff, but partnerships lead to slight rebound

Makers of software that runs large companies saw their share prices plunge as investors worried that AI systems could undermine their businesses. This week, their stocks rebounded somewhat as Anthropic partnered with some of the same companies.

February 27, 20263 min read
Faster Reasoning at the Edge: Liquid AI’s small reasoning model mixes attention with convolutional layers for efficiency
Large Language Models (LLMs)

Faster Reasoning at the Edge: Liquid AI’s small reasoning model mixes attention with convolutional layers for efficiency

Reasoning models in the 1 to 2 billion-parameter range typically require more than 1 gigabyte of RAM to run. Liquid AI released one that runs in less than 900 megabytes, and does it with exceptional speed and efficiency.

February 20, 20263 min read
GLM-5 Scales Up: Z.ai’s updated model boasts top open-weights Intelligence Index score
Large Language Models (LLMs)

GLM-5 Scales Up: Z.ai’s updated model boasts top open-weights Intelligence Index score

Z.ai more than doubled the size of its flagship large language model to deliver outstanding performance among open-weights competitors.

February 20, 20262 min read
xAI Blasts Off: SpaceX acquires xAI, announces plans for data centers In space
Large Language Models (LLMs)

xAI Blasts Off: SpaceX acquires xAI, announces plans for data centers In space

Elon Musk’s SpaceX acquired xAI, opening the door to richer financing of the merged entity’s AI research, a tighter focus on space applications of AI, and — if Musk’s dreams are realized — solar-powered data centers in space.

February 13, 20263 min read
Claude Opus 4.6 Reasons More Over Harder Problems: Anthropic updates flagship model, places first on Intelligence Index
Large Language Models (LLMs)

Claude Opus 4.6 Reasons More Over Harder Problems: Anthropic updates flagship model, places first on Intelligence Index

Anthropic updated its flagship large language model to handle longer, more complex agentic tasks.

February 13, 20263 min read
AI Giants Share Wikipedia’s Costs: Wikimedia Foundation strikes deals with Amazon, Meta, Microsoft, Mistral AI, and Perplexity
Large Language Models (LLMs)

AI Giants Share Wikipedia’s Costs: Wikimedia Foundation strikes deals with Amazon, Meta, Microsoft, Mistral AI, and Perplexity

On its 25th anniversary, Wikipedia celebrated with high-profile deals to make its data easier for AI companies to train their models in exchange for financial support.

February 6, 20263 min read
Training For Engagement Can Degrade Alignment: “Moloch’s Bargain” shows fine-tuning can affect social values
Large Language Models (LLMs)

Training For Engagement Can Degrade Alignment: “Moloch’s Bargain” shows fine-tuning can affect social values

Individuals and organizations increasingly use large language models to produce media that helps them compete for attention. Does fine-tuning LLMs to encourage engagement, purchases, or votes affect their alignment with social values? Researchers found that it does.

January 30, 20263 min read
Artificial Analysis Revamps Intelligence Index: Independent AI testing authority turns from saturated knowledge benchmarks to harder business tests
Large Language Models (LLMs)

Artificial Analysis Revamps Intelligence Index: Independent AI testing authority turns from saturated knowledge benchmarks to harder business tests

Artificial Analysis, which tests AI systems, updated the component evaluations in its Intelligence Index to better reflect large language models’ performance in real-world use cases.

January 30, 20263 min read
Apple’s Foundation Models Will Be Gemini: Apple announced a partnership with Google to power Siri and other AI features
Large Language Models (LLMs)

Apple’s Foundation Models Will Be Gemini: Apple announced a partnership with Google to power Siri and other AI features

Apple cut a multi-year deal with Google to use Gemini models as the basis of AI models that reside on Apple devices.

January 23, 20263 min read
ChatGPT Shows Ads: OpenAI tests advertisements for U.S. chatbot users in free and lower-cost tiers
Large Language Models (LLMs)

ChatGPT Shows Ads: OpenAI tests advertisements for U.S. chatbot users in free and lower-cost tiers

AI has a new revenue stream, and it looks a lot like old web banner ads.

January 23, 20263 min read
Retrieval Faces Hard Limits: Google and Johns Hopkins researchers show embedding models can’t search unlimited documents
Large Language Models (LLMs)

Retrieval Faces Hard Limits: Google and Johns Hopkins researchers show embedding models can’t search unlimited documents

Can your retriever find all the relevant documents for any query your users might enter? Maybe not, research shows.

January 16, 20264 min read
More Affordable Reasoning: Canadian researchers find capping context helps models better retrieve data
Large Language Models (LLMs)

More Affordable Reasoning: Canadian researchers find capping context helps models better retrieve data

One way to improve a reasoning model’s performance is to let it produce a longer chain of thought. However, attending to ever-longer contexts can become expensive, and making that attention more efficient requires changes to a model’s architecture.

January 9, 20262 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox