Cursor Fits Its Model to Its Agent: Composer 2.5 for Cursor rivals GPT-5.5's coding abilities at lower price
Large Language Models (LLMs)

Cursor Fits Its Model to Its Agent: Composer 2.5 for Cursor rivals GPT-5.5's coding abilities at lower price

Cursor’s latest software engineering model rivals the performance of leading competitors like Claude Opus 4.7 and GPT 5.5 for a fraction of the price.

June 12, 20264 min read
Behold Mythos!: Anthropic released Claude Mythos 5 and Claude Fable 5, a public version with safeguards
Large Language Models (LLMs)

Behold Mythos!: Anthropic released Claude Mythos 5 and Claude Fable 5, a public version with safeguards

After months of headlines that teased a large language model with extraordinary capabilities, Anthropic launched Claude Mythos 5, which can crack software previously believed to be secure, and Claude Fable 5, a version for general use that limits what users can do in an unprecedented way.

June 12, 20265 min read
Fine-Tuning LLMs to Expand on Summaries Unearths Pretraining Texts: Fine-Tuning can strip models of copyright alignment guidelines
Large Language Models (LLMs)

Fine-Tuning LLMs to Expand on Summaries Unearths Pretraining Texts: Fine-Tuning can strip models of copyright alignment guidelines

Fine-tuning large language models on a seemingly benign task that would be useful to writers — expanding plot summaries into paragraphs of polished fiction — causes them to regurgitate substantial portions of books on which they were pretrained.

June 5, 20263 min read
Inside the Gray Market for LLM Access: Middlemen package extra tokens, hijack IDs to resell, distill models
Large Language Models (LLMs)

Inside the Gray Market for LLM Access: Middlemen package extra tokens, hijack IDs to resell, distill models

An ecosystem of API proxy servers enables AI developers in China to access top U.S. models at deeply discounted prices.

June 5, 20263 min read
Qwen3.7-Max Adds Speed and Power: Alibaba's latest proprietary model challenges U.S. rivals
Large Language Models (LLMs)

Qwen3.7-Max Adds Speed and Power: Alibaba's latest proprietary model challenges U.S. rivals

Alibaba updated its flagship large language model for long-running agentic work, pushing it into the top rank among LLMs built in China.

June 5, 20263 min read
Cybersecurity Alarms Grow Louder: Google study shows LLM-generated malware is getting harder to track and stop
Large Language Models (LLMs)

Cybersecurity Alarms Grow Louder: Google study shows LLM-generated malware is getting harder to track and stop

An AI-generated script to bypass two-factor authentication signals a dawning era of industrial-scale cyberattacks, according to a Google report.

May 22, 20263 min read
Hermes Agent Challenges OpenClaw: OpenClaw created a class of personal agents; upstart Hermes Agent is outworking it
Large Language Models (LLMs)

Hermes Agent Challenges OpenClaw: OpenClaw created a class of personal agents; upstart Hermes Agent is outworking it

OpenClaw, the immensely popular AI agent, has fast-rising competition.

May 22, 20264 min read
Assistants That Assist Consistently: Large language models can drift drift from helpful personas to harmful ones, but new research aims to stabilize them
Large Language Models (LLMs)

Assistants That Assist Consistently: Large language models can drift drift from helpful personas to harmful ones, but new research aims to stabilize them

Typically, large language models are trained to act as helpful, harmless, honest assistants. However, during long or emotionally charged conversations, traits can emerge that are less beneficial. Researchers devised a way to steady the assistant personas of LLMs.

April 24, 20263 min read
GLM-5.1 Aims for Long-Running Tasks: Z.ai’s GLM 5.1 evaluates interim results and may change its approach hundreds of times before it delivers final output
Large Language Models (LLMs)

GLM-5.1 Aims for Long-Running Tasks: Z.ai’s GLM 5.1 evaluates interim results and may change its approach hundreds of times before it delivers final output

Z.ai updated its flagship open-weights large language model to work autonomously on single tasks for up to eight hours.

April 24, 20263 min read
Simulating Diverse Human Cohorts: Persona generation simulates human characters across a controllable range of points of view
Large Language Models (LLMs)

Simulating Diverse Human Cohorts: Persona generation simulates human characters across a controllable range of points of view

If you want to understand how the public will respond to your offerings, large language models can simulate users who answer questions about capabilities, features, promotions, or prices.

April 17, 20263 min read
Claude Mythos Preview Raises Security Worries: Why Claude’s advanced Mythos Preview model will be limited-release-only
Large Language Models (LLMs)

Claude Mythos Preview Raises Security Worries: Why Claude’s advanced Mythos Preview model will be limited-release-only

Anthropic took unusual steps to prepare the world for a forthcoming large language model that it said poses extraordinary risks to cybersecurity.

April 10, 20264 min read
Learning Long Context at Inference: Test-Time Training End-to-End (TTT-E2E) retrains model weights to handle long inputs
Large Language Models (LLMs)

Learning Long Context at Inference: Test-Time Training End-to-End (TTT-E2E) retrains model weights to handle long inputs

Large language models typically become less accurate and slower when they process longer contexts, but researchers enabled an LLM to keep accuracy stable and inference time constant as its context grew.

April 3, 20264 min read
Context As An External Variable: Recursive Language Models offer path to aramatically expand beyond the context window
Large Language Models (LLMs)

Context As An External Variable: Recursive Language Models offer path to aramatically expand beyond the context window

When processing long contexts, large language models often lose track of details or devolve into nonsense. Researchers reduced these effects by managing context externally.

March 27, 20263 min read
Open-Source Speed Demon: Nvidia’s open Nemotron 3 Super 120B-A12B model sets new paces in its class
Large Language Models (LLMs)

Open-Source Speed Demon: Nvidia’s open Nemotron 3 Super 120B-A12B model sets new paces in its class

Nvidia, the dominant supplier of AI chips, released a competitive open-source large language model whose speed tops its size class — the first open-weights leader to come from the United States since last year, when Meta delivered Llama 4.

March 27, 20264 min read
AI on Mobile Skyrockets: State of Mobile 2026 Report shows AI chatbot, search, and assistant growth outpaces gaming, social, and more
Large Language Models (LLMs)

AI on Mobile Skyrockets: State of Mobile 2026 Report shows AI chatbot, search, and assistant growth outpaces gaming, social, and more

Downloads of mobile AI apps and resulting revenue are surging.

March 16, 20262 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox