Reinforcement Learning Heats Up: How DeepSeek-R1 and Kimi k1.5 use reinforcement learning to improve reasoning
Large Language Models (LLMs)

Reinforcement Learning Heats Up: How DeepSeek-R1 and Kimi k1.5 use reinforcement learning to improve reasoning

Reinforcement learning is emerging as an avenue for building large language models with advanced reasoning capabilities.

January 29, 20252 min read
DeepSeek Sharpens Its Reasoning: DeepSeek-R1, an affordable rival to OpenAI’s o1
Large Language Models (LLMs)

DeepSeek Sharpens Its Reasoning: DeepSeek-R1, an affordable rival to OpenAI’s o1

A new open model rivals OpenAI’s o1, and it’s free to use or modify.

January 22, 20254 min read
DeepSeek Sharpens Its Reasoning: DeepSeek-R1, an affordable rival to OpenAI’s o1
Large Language Models (LLMs)

DeepSeek Sharpens Its Reasoning: DeepSeek-R1, an affordable rival to OpenAI’s o1

A new open model rivals OpenAI’s o1, and it’s free to use or modify.

January 22, 20254 min read
DeepSeek Ups the Open Weights Ante: DeepSeek-V3 redefines LLM performance and cost efficiency
Large Language Models (LLMs)

DeepSeek Ups the Open Weights Ante: DeepSeek-V3 redefines LLM performance and cost efficiency

A new model from Hangzhou upstart DeepSeek delivers outstanding performance and may change the equation for training costs.

January 15, 20253 min read
DeepSeek Ups the Open Weights Ante: DeepSeek-V3 redefines LLM performance and cost efficiency
Large Language Models (LLMs)

DeepSeek Ups the Open Weights Ante: DeepSeek-V3 redefines LLM performance and cost efficiency

A new model from Hangzhou upstart DeepSeek delivers outstanding performance and may change the equation for training costs.

January 15, 20253 min read
Better Performance From Merged Models: Localize-and-Stitch improves methods for merging and fine-tuning multiple models
Large Language Models (LLMs)

Better Performance From Merged Models: Localize-and-Stitch improves methods for merging and fine-tuning multiple models

Merging multiple fine-tuned models is a less expensive alternative to hosting multiple specialized models. But, while model merging can deliver higher average performance across several tasks, it often results in lower performance on specific tasks. New work addresses this issue.

January 8, 20253 min read
Better Performance From Merged Models: Localize-and-Stitch improves methods for merging and fine-tuning multiple models
Large Language Models (LLMs)

Better Performance From Merged Models: Localize-and-Stitch improves methods for merging and fine-tuning multiple models

Merging multiple fine-tuned models is a less expensive alternative to hosting multiple specialized models. But, while model merging can deliver higher average performance across several tasks, it often results in lower performance on specific tasks. New work addresses this issue.

January 8, 20253 min read
Massively More Training Text: Harvard unveils a million-book corpus for AI training
Large Language Models (LLMs)

Massively More Training Text: Harvard unveils a million-book corpus for AI training

Harvard University amassed a huge new text corpus for training machine learning models.

January 8, 20252 min read
Massively More Training Text: Harvard unveils a million-book corpus for AI training
Large Language Models (LLMs)

Massively More Training Text: Harvard unveils a million-book corpus for AI training

Harvard University amassed a huge new text corpus for training machine learning models.

January 8, 20252 min read
Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors
Large Language Models (LLMs)

Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

January 8, 20256 min read
Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors
Large Language Models (LLMs)

Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

January 8, 20256 min read
What LLM Users Want: Anthropic reveals how users interact with Claude 3.5
Large Language Models (LLMs)

What LLM Users Want: Anthropic reveals how users interact with Claude 3.5

Anthropic analyzed 1 million anonymized conversations between users and Claude 3.5 Sonnet. The study found that most people used the model for software development and also revealed malfunctions and jailbreaks.

January 8, 20252 min read
What LLM Users Want: Anthropic reveals how users interact with Claude 3.5
Large Language Models (LLMs)

What LLM Users Want: Anthropic reveals how users interact with Claude 3.5

Anthropic analyzed 1 million anonymized conversations between users and Claude 3.5 Sonnet. The study found that most people used the model for software development and also revealed malfunctions and jailbreaks.

January 8, 20252 min read
Joseph Gonzalez: General intelligence
Large Language Models (LLMs)

Joseph Gonzalez: General intelligence

In 2025, I expect progress in training foundation models to slow down as we hit scaling limits and inference costs continue to rise.

January 1, 20253 min read
Joseph Gonzalez: General intelligence
Large Language Models (LLMs)

Joseph Gonzalez: General intelligence

In 2025, I expect progress in training foundation models to slow down as we hit scaling limits and inference costs continue to rise.

January 1, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox