Good Models, Bad Choices: Anthropic made LLMs choose between failing and misbehaving, and they blackmailed executives.
Large Language Models (LLMs)

Good Models, Bad Choices: Anthropic made LLMs choose between failing and misbehaving, and they blackmailed executives.

Top large language models, under experimental conditions that pressed them to choose between abandoning their prompted mission and misbehaving, resorted to harmful behavior, researchers found.

July 9, 20254 min read
Reasoning for No Reason: Anthropic finds chain-of-thought reasoning traces may omit key influences
Large Language Models (LLMs)

Reasoning for No Reason: Anthropic finds chain-of-thought reasoning traces may omit key influences

Does a reasoning model’s chain of thought explain how it arrived at its output? Researchers found that often it doesn’t.

July 2, 20252 min read
Low Precision, High Performance: Researchers at Microsoft and Tsinghua researchers propose 1.58-bit AI model that rivals full-precision competitors
Large Language Models (LLMs)

Low Precision, High Performance: Researchers at Microsoft and Tsinghua researchers propose 1.58-bit AI model that rivals full-precision competitors

Reducing the number of bits used to represent each parameter in a neural network from, say, 16 bits to 8 bits shrinks the network’s size and boosts its speed. Researchers took this approach to an extreme: They built a competitive large language model whose weights are limited to three values.

June 25, 20253 min read
Apple Sharpens Its GenAI Profile: Apple updates its on-device and cloud AI models, introduces a new developer API
Large Language Models (LLMs)

Apple Sharpens Its GenAI Profile: Apple updates its on-device and cloud AI models, introduces a new developer API

Apple revamped two vision-language models in a bid to catch up with fast-moving competitors.

June 19, 20254 min read
LLM Rights Historical Wrongs: Stanford and Princeton researchers fine-tune a language model to identify racial discrimination in property
Large Language Models (LLMs)

LLM Rights Historical Wrongs: Stanford and Princeton researchers fine-tune a language model to identify racial discrimination in property

In Northern California, old property deeds may still include racial clauses: language, made illegal decades ago, that was designed to ban people of color from owning or living in certain homes.

June 19, 20252 min read
More Reasoning for Harder Problems: OpenAI debuts o3-pro, an updated reasoning model that applies more tokens at inference
Large Language Models (LLMs)

More Reasoning for Harder Problems: OpenAI debuts o3-pro, an updated reasoning model that applies more tokens at inference

OpenAI launched o3-pro, a more capable version of its most advanced reasoning vision-language model.

June 19, 20252 min read
Better Video, Fewer Tokens: STORM Processes Fewer Tokens And Still Beats GPT-4o On Video Understanding Benchmarks
Large Language Models (LLMs)

Better Video, Fewer Tokens: STORM Processes Fewer Tokens And Still Beats GPT-4o On Video Understanding Benchmarks

Researchers reduced the number of tokens needed to represent video frames to be fed to a transformer.

June 11, 20252 min read
AI Market Trends in Charts and Graphs: Venture Capitalist Mary Meeker Revives Her Trend Reports With a Deep Dive Into the AI Boom
Large Language Models (LLMs)

AI Market Trends in Charts and Graphs: Venture Capitalist Mary Meeker Revives Her Trend Reports With a Deep Dive Into the AI Boom

Renowned investment analyst Mary Meeker is back with a report on the AI market, six years after publishing her last survey of the internet.

June 11, 20253 min read
Next-Level DeepSeek-R1: DeepSeek-R1’s update leads all open models and brings it up to date with the latest from Google and OpenAI
Large Language Models (LLMs)

Next-Level DeepSeek-R1: DeepSeek-R1’s update leads all open models and brings it up to date with the latest from Google and OpenAI

DeepSeek updated its groundbreaking DeepSeek-R1 large language model to strike another blow for open-weights performance.

June 4, 20253 min read
How DeepSeek Did It: Researchers describe training methods and hardware choices for DeepSeek’s V3 and R1 models
Large Language Models (LLMs)

How DeepSeek Did It: Researchers describe training methods and hardware choices for DeepSeek’s V3 and R1 models

DeepSeek made headlines late last year, when it built a state-of-the-art, open-weights large language model at a cost far lower than usual. The upstart developer shared new details about its method.

May 28, 20253 min read
Claude 4 Advances Code Generation: Anthropic debuts new Claude 4 Sonnet and Claude 4 Opus models, featuring top benchmarks in coding
Large Language Models (LLMs)

Claude 4 Advances Code Generation: Anthropic debuts new Claude 4 Sonnet and Claude 4 Opus models, featuring top benchmarks in coding

Anthropic continued its tradition of building AI models that raise the bar in coding tasks.

May 28, 20253 min read
4-Bit Efficiency, 16-Bit Accuracy: Microsoft researchers show that heavily quantized versions of Llama can perform as well as near-full-precision
Large Language Models (LLMs)

4-Bit Efficiency, 16-Bit Accuracy: Microsoft researchers show that heavily quantized versions of Llama can perform as well as near-full-precision

Using an 8-bit number format like FP8 during training saves computation compared to 16- or 32-bit formats, but it can yield less-accurate results. Researchers trained models using 4-bit numbers without sacrificing accuracy.

May 21, 20252 min read
Your Robot Dev Team: OpenAI introduces Codex, a multi-agent cloud-based software engineering tool in ChatGPT
Large Language Models (LLMs)

Your Robot Dev Team: OpenAI introduces Codex, a multi-agent cloud-based software engineering tool in ChatGPT

OpenAI launched an agentic software-development system.

May 21, 20253 min read
Memory Layers for More-Factual Output: Meta researchers build Llama-style models that recall details without needing more computing resources
Large Language Models (LLMs)

Memory Layers for More-Factual Output: Meta researchers build Llama-style models that recall details without needing more computing resources

Improving a large language model’s factual accuracy typically requires making it bigger, which in turn, involves more computation. Researchers devised an architecture that enables models to recall relevant details without significantly increasing the amount of computation required.

May 14, 20253 min read
Open, Compact Code Generator: DeepCoder-14B-Preview further fine-tunes reasoning models for coding
Large Language Models (LLMs)

Open, Compact Code Generator: DeepCoder-14B-Preview further fine-tunes reasoning models for coding

An open-source code generator performs comparably to the reasoning models DeepSeek-R1 and OpenAI o1 with a much smaller model.

May 14, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox