DeepSeek-R1 Uncensored: Perplexity launches uncensored version of DeepSeek-R1
Large Language Models (LLMs)

DeepSeek-R1 Uncensored: Perplexity launches uncensored version of DeepSeek-R1

Large language models built by developers in China may, in some applications, be less useful outside that country because they avoid topics its government deems politically sensitive. A developer fine-tuned DeepSeek-R1 to widen its scope without degrading its overall performance.

March 12, 20252 min read
Compact Reasoning: QwQ-32B challenges DeepSeek-R1 and other larger reasoning models
Large Language Models (LLMs)

Compact Reasoning: QwQ-32B challenges DeepSeek-R1 and other larger reasoning models

Most models that have learned to reason via reinforcement learning were huge models. A much smaller model now competes with them.

March 12, 20252 min read
Compact Reasoning: QwQ-32B challenges DeepSeek-R1 and other larger reasoning models
Large Language Models (LLMs)

Compact Reasoning: QwQ-32B challenges DeepSeek-R1 and other larger reasoning models

Most models that have learned to reason via reinforcement learning were huge models. A much smaller model now competes with them.

March 12, 20252 min read
Budget for Reasoning to the Token: Claude 3.7 Sonnet adds extended thinking mode
Large Language Models (LLMs)

Budget for Reasoning to the Token: Claude 3.7 Sonnet adds extended thinking mode

Anthropic’s Claude 3.7 Sonnet implements a hybrid reasoning approach that lets users decide how much thinking they want the model to do before it renders a response.

March 5, 20253 min read
Budget for Reasoning to the Token: Claude 3.7 Sonnet adds extended thinking mode
Large Language Models (LLMs)

Budget for Reasoning to the Token: Claude 3.7 Sonnet adds extended thinking mode

Anthropic’s Claude 3.7 Sonnet implements a hybrid reasoning approach that lets users decide how much thinking they want the model to do before it renders a response.

March 5, 20253 min read
OpenAI’s GPT-4.5 Goes Big: OpenAI releases GPT-4.5, its most powerful non-reasoning model and maybe its last
Large Language Models (LLMs)

OpenAI’s GPT-4.5 Goes Big: OpenAI releases GPT-4.5, its most powerful non-reasoning model and maybe its last

OpenAI launched GPT-4.5, which may be its last non-reasoning model.

March 5, 20253 min read
OpenAI’s GPT-4.5 Goes Big: OpenAI releases GPT-4.5, its most powerful non-reasoning model and maybe its last
Large Language Models (LLMs)

OpenAI’s GPT-4.5 Goes Big: OpenAI releases GPT-4.5, its most powerful non-reasoning model and maybe its last

OpenAI launched GPT-4.5, which may be its last non-reasoning model.

March 5, 20253 min read
Text Generation by Diffusion: Mercury Coder uses diffusion to generate text
Large Language Models (LLMs)

Text Generation by Diffusion: Mercury Coder uses diffusion to generate text

Typical large language models are autoregressive, predicting the next token, one at a time, from left to right. A new model hones all text tokens at once.

March 5, 20252 min read
Text Generation by Diffusion: Mercury Coder uses diffusion to generate text
Large Language Models (LLMs)

Text Generation by Diffusion: Mercury Coder uses diffusion to generate text

Typical large language models are autoregressive, predicting the next token, one at a time, from left to right. A new model hones all text tokens at once.

March 5, 20252 min read
Reasoning in Vectors, Not Text: Meta introduces Chain of Continuous Thought (Coconut) to improve next-token prediction
Large Language Models (LLMs)

Reasoning in Vectors, Not Text: Meta introduces Chain of Continuous Thought (Coconut) to improve next-token prediction

Although large language models can improve their performance by generating a chain of thought (CoT) — intermediate text tokens that break down the process of responding to a prompt into a series of steps.

February 26, 20253 min read
Reasoning in Vectors, Not Text: Meta introduces Chain of Continuous Thought (Coconut) to improve next-token prediction
Large Language Models (LLMs)

Reasoning in Vectors, Not Text: Meta introduces Chain of Continuous Thought (Coconut) to improve next-token prediction

Although large language models can improve their performance by generating a chain of thought (CoT) — intermediate text tokens that break down the process of responding to a prompt into a series of steps.

February 26, 20253 min read
Musk Complicates OpenAI’s Plan: Elon Musk’s $97.4B bid for OpenAI rejected, fueling AI power struggle
Large Language Models (LLMs)

Musk Complicates OpenAI’s Plan: Elon Musk’s $97.4B bid for OpenAI rejected, fueling AI power struggle

Elon Musk and a group of investors made an unsolicited bid to buy the assets of the nonprofit that controls OpenAI, complicating the AI powerhouse’s future plans.

February 19, 20253 min read
Musk Complicates OpenAI’s Plan: Elon Musk’s $97.4B bid for OpenAI rejected, fueling AI power struggle
Large Language Models (LLMs)

Musk Complicates OpenAI’s Plan: Elon Musk’s $97.4B bid for OpenAI rejected, fueling AI power struggle

Elon Musk and a group of investors made an unsolicited bid to buy the assets of the nonprofit that controls OpenAI, complicating the AI powerhouse’s future plans.

February 19, 20253 min read
Mobile Apps to Order: Replit’s agent-powered mobile app expands to full app development
Large Language Models (LLMs)

Mobile Apps to Order: Replit’s agent-powered mobile app expands to full app development

Replit, an AI-driven integrated development environment, updated its mobile app to generate further mobile apps to order.

February 19, 20252 min read
Mobile Apps to Order: Replit’s agent-powered mobile app expands to full app development
Large Language Models (LLMs)

Mobile Apps to Order: Replit’s agent-powered mobile app expands to full app development

Replit, an AI-driven integrated development environment, updated its mobile app to generate further mobile apps to order.

February 19, 20252 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox