Reasoning Models With Recipes: Microsoft unveils training details for Phi-4-reasoning, Phi-4-reasoning-plus, and Phi-4-mini-reasoning
Large Language Models (LLMs)

Reasoning Models With Recipes: Microsoft unveils training details for Phi-4-reasoning, Phi-4-reasoning-plus, and Phi-4-mini-reasoning

Microsoft published its latest recipe for training reasoning models, substantially expanding what is still a fairly small base of public knowledge.

May 14, 20253 min read
One Weird Trick for Better Reasoning: Researchers fine-tune LLM for reasoning with only 1,000 examples
Large Language Models (LLMs)

One Weird Trick for Better Reasoning: Researchers fine-tune LLM for reasoning with only 1,000 examples

Researchers showed that supervised fine-tuning on as few as 1,000 examples can enable a pretrained large language model to reason — and a clever gambit can boost its performance to rival that of top reasoning models.

May 7, 20253 min read
Qwen3 Takes On DeepSeek-R1: Alibaba releases the Qwen3 family of open LLMs with optional reasoning
Large Language Models (LLMs)

Qwen3 Takes On DeepSeek-R1: Alibaba releases the Qwen3 family of open LLMs with optional reasoning

Alibaba’s new model family may unseat DeepSeek-R1’s four-month reign as the top open-weights large language model.

May 7, 20253 min read
Inferring Customer Preferences: LLMs boost shopping recommendations by decoding what users want
Large Language Models (LLMs)

Inferring Customer Preferences: LLMs boost shopping recommendations by decoding what users want

Large language models can improve systems that recommend items to purchase by inferring customer preferences.

April 30, 20252 min read
OpenAI Launches Cost-Effective Alternatives: OpenAI replaces GPT-4.5 with GPT-4.1 Family, plus o3 and o4-mini, new models focused on reasoning and coding
Large Language Models (LLMs)

OpenAI Launches Cost-Effective Alternatives: OpenAI replaces GPT-4.5 with GPT-4.1 Family, plus o3 and o4-mini, new models focused on reasoning and coding

OpenAI refreshed its roster of models and scheduled the largest, most costly one for removal.

April 23, 20253 min read
Toward LLMs That Understand Misspellings: New byte-based model beats Llama 3 on spelling, noise, and translation
Large Language Models (LLMs)

Toward LLMs That Understand Misspellings: New byte-based model beats Llama 3 on spelling, noise, and translation

Researchers built a model that’s more robust to noisy inputs like misspellings, smarter about character-level information like the number of R's in strawberry, and potentially better able to understand unfamiliar languages that might share groups of letters with familiar languages.

April 16, 20253 min read
The Fall and Rise of Sam Altman: Inside Sam Altman’s brief ouster from OpenAI
Large Language Models (LLMs)

The Fall and Rise of Sam Altman: Inside Sam Altman’s brief ouster from OpenAI

A behind-the-scenes account provides new details about the abrupt firing and reinstatement of OpenAI CEO Sam Altman in November 2023.

April 16, 20252 min read
Open Standard for Tool Use and Data Access Gains Momentum: OpenAI adopts Model Context Protocol to boost LLM tool integration
Large Language Models (LLMs)

Open Standard for Tool Use and Data Access Gains Momentum: OpenAI adopts Model Context Protocol to boost LLM tool integration

OpenAI embraced Model Context Protocol, providing powerful support for an open standard that connects large language models to tools and data.

April 16, 20252 min read
Google Unveils Gemini 2.5: Google’s Gemini 2.5 Pro Experimental outperforms top AI models
Large Language Models (LLMs)

Google Unveils Gemini 2.5: Google’s Gemini 2.5 Pro Experimental outperforms top AI models

Google’s new flagship model raised the state of the art in a variety of subjective and objective tests.

April 16, 20252 min read
Llama’s Mixture of Vision-Language Experts: Meta releases Llama 4 models, claims edge over AI competitors
Large Language Models (LLMs)

Llama’s Mixture of Vision-Language Experts: Meta releases Llama 4 models, claims edge over AI competitors

Meta updated its popular open-weights models, claiming performance superior to closed competitors in three size classes.

April 9, 20253 min read
LLM Support for Tutors: GPT-4 boosts remote tutors’ performance in real time, study finds
Large Language Models (LLMs)

LLM Support for Tutors: GPT-4 boosts remote tutors’ performance in real time, study finds

Students benefit from tutoring, but training tutors is expensive. A study shows that large language models can boost tutors’ effectiveness in real time.

March 26, 20252 min read
Vision-Language, Compact and Open: Google releases Gemma 3 vision-language models with open weights
Large Language Models (LLMs)

Vision-Language, Compact and Open: Google releases Gemma 3 vision-language models with open weights

Google updated its open-weights family of large language models to include versions that handle image and video inputs.

March 26, 20253 min read
Some AI-Generated Works Are Copyrightable: U.S. Copyright Office says that no new laws are needed for AI-generated works
Large Language Models (LLMs)

Some AI-Generated Works Are Copyrightable: U.S. Copyright Office says that no new laws are needed for AI-generated works

The United States Copyright Office determined that existing laws are sufficient to decide whether a given AI-generated work is protected by copyright, making additional legislation unnecessary.

March 19, 20252 min read
Some AI-Generated Works Are Copyrightable: U.S. Copyright Office says that no new laws are needed for AI-generated works
Large Language Models (LLMs)

Some AI-Generated Works Are Copyrightable: U.S. Copyright Office says that no new laws are needed for AI-generated works

The United States Copyright Office determined that existing laws are sufficient to decide whether a given AI-generated work is protected by copyright, making additional legislation unnecessary.

March 19, 20252 min read
DeepSeek-R1 Uncensored: Perplexity launches uncensored version of DeepSeek-R1
Large Language Models (LLMs)

DeepSeek-R1 Uncensored: Perplexity launches uncensored version of DeepSeek-R1

Large language models built by developers in China may, in some applications, be less useful outside that country because they avoid topics its government deems politically sensitive. A developer fine-tuned DeepSeek-R1 to widen its scope without degrading its overall performance.

March 12, 20252 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox