Large Language Models (LLMs)

353 Posts

Web Retrieval Flusters LLMs: AI agents searching online can struggle to retrieve correct info, researchers find
Large Language Models (LLMs)

Web Retrieval Flusters LLMs: AI agents searching online can struggle to retrieve correct info, researchers find

Large language models often are called upon to gather news. In this task, researchers found, their ability to find relevant reports is the weakest link.

July 24, 20263 min read
Meta Sparks A Price War: Muse Spark 1.1 makes a jump in intelligence at a lower price than peers
Large Language Models (LLMs)

Meta Sparks A Price War: Muse Spark 1.1 makes a jump in intelligence at a lower price than peers

With Llama, Meta marked itself as an open alternative to OpenAI. With its new closed models, Meta now positions itself as a low-cost, high-value competitor.

July 24, 20264 min read
Measuring Models’ Manipulation: Researchers at MIT and Carnegie Mellon built Puppet to gauge models' influence on users' beliefs
Large Language Models (LLMs)

Measuring Models’ Manipulation: Researchers at MIT and Carnegie Mellon built Puppet to gauge models' influence on users' beliefs

Providers of large language models stand to benefit by building models that spur user engagement, but users may bear a cost in undue influence on their world views.

July 17, 20264 min read
One Model Talks, Another One Thinks: GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) on the backend
Large Language Models (LLMs)

One Model Talks, Another One Thinks: GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) on the backend

ChatGPT’s voice mode now listens and speaks at the same time, passing harder questions posed to the conversational model to a reasoning model in the background.

July 17, 20265 min read
DeepSeek’s DSpark Gains Velocity: DeepSeek open sourced a speculative decoding module that speeds up text generation without losing accuracy
Large Language Models (LLMs)

DeepSeek’s DSpark Gains Velocity: DeepSeek open sourced a speculative decoding module that speeds up text generation without losing accuracy

DeepSeek built a speculative decoding module that speeds up its production models’ text generation by more than 50 percent without sacrificing accuracy, then made its technique open source.

July 10, 20265 min read
Microsoft Strikes Out on Its Own: Microsoft revealed MAI-Thinking-1, a Claude Sonnet 4.6-sized reasoning model developed without distillation
Large Language Models (LLMs)

Microsoft Strikes Out on Its Own: Microsoft revealed MAI-Thinking-1, a Claude Sonnet 4.6-sized reasoning model developed without distillation

Microsoft, once OpenAI’s exclusive partner and still a major reseller of other companies’ AI models, built its own reasoning model from scratch.

July 3, 20263 min read
Fugu Blends Models Task by Task: Sakana debuted dedicated orchestrator models, Fugu and Fugu-Ultra, that spawn Claude, Gemini, and GPT agents
Large Language Models (LLMs)

Fugu Blends Models Task by Task: Sakana debuted dedicated orchestrator models, Fugu and Fugu-Ultra, that spawn Claude, Gemini, and GPT agents

Models that orchestrate the activities of other models and agents achieved state-of-the-art performance on a variety of benchmarks, outperforming the best individual models working alone.

July 3, 20264 min read
GPT-5.6 Lands in Limbo: OpenAI previewed three GPT-5.6 Models (Sol, Terra, and Luna), wider release coming soon
Large Language Models (LLMs)

GPT-5.6 Lands in Limbo: OpenAI previewed three GPT-5.6 Models (Sol, Terra, and Luna), wider release coming soon

OpenAI announced a preview of its GPT-5.6 family, including a top-tier model comparable to Claude 5 Mythos — but so far it’s available only to users that are selected by the U.S. government.

July 3, 20265 min read
Biological Molecules as Language: ESMFold2 approaches AlphaFold 3 performance but with an open, Transformer-based architecture
Large Language Models (LLMs)

Biological Molecules as Language: ESMFold2 approaches AlphaFold 3 performance but with an open, Transformer-based architecture

Google’s AlphaFold models pioneered the task of finding the shapes of biologically active molecules, opening new pathways for drug development.

June 26, 20264 min read
Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud
Large Language Models (LLMs)

Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud

The third generation of Apple Foundation Models — fruit of Apple’s collaboration with Google — introduces a variation on the mixture-of-experts architecture that runs on local devices. 

June 26, 20263 min read
Top Agentic Performance, Low Cost: GLM-5.2, designed for coding and long-running agentic jobs, now the top open model
Large Language Models (LLMs)

Top Agentic Performance, Low Cost: GLM-5.2, designed for coding and long-running agentic jobs, now the top open model

Z.ai released an open-weights model that rivals proprietary leaders for autonomous agentic tasks.

June 26, 20264 min read
Reinforcement Learning With Hints: Privileged On-Policy Exploration (POPE) trains models to expand on partial solutions
Large Language Models (LLMs)

Reinforcement Learning With Hints: Privileged On-Policy Exploration (POPE) trains models to expand on partial solutions

Reinforcement learning can’t train a model to solve a difficult problem if the model doesn’t discover all the right steps.

June 19, 20263 min read
Nvidia’s Nemotron Goes Big: Nvidia Nemotron 3 Ultra bets on speed and openness to win customers
Large Language Models (LLMs)

Nvidia’s Nemotron Goes Big: Nvidia Nemotron 3 Ultra bets on speed and openness to win customers

Nvidia’s largest-yet model is among the best-performing from a developer based in the U.S. and among the most open developed by anyone.

June 19, 20264 min read
Agentic Tests Beyond the Bug Hunt: DeepSWE, ProgramBench, and ITBench-AA push agents harder than SWE-bench
Large Language Models (LLMs)

Agentic Tests Beyond the Bug Hunt: DeepSWE, ProgramBench, and ITBench-AA push agents harder than SWE-bench

SWE-bench, a family of benchmarks that focuses on an LLM’s ability to fix software bugs, is giving way to new tests that evaluate agent software-engineering performance in more challenging ways.

June 19, 20264 min read
State Media Influences LLM Responses: Significant portions of AI training material reflect national propaganda
Large Language Models (LLMs)

State Media Influences LLM Responses: Significant portions of AI training material reflect national propaganda

Popular large language models have adopted the biases of governments that control the free flow of information, particularly when those models generate output in the languages of countries where such governments are in power, researchers found.

June 12, 20264 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox