
Reinforcement Learning Heats Up: How DeepSeek-R1 and Kimi k1.5 use reinforcement learning to improve reasoning
Reinforcement learning is emerging as an avenue for building large language models with advanced reasoning capabilities.

Reinforcement learning is emerging as an avenue for building large language models with advanced reasoning capabilities.

A new open model rivals OpenAI’s o1, and it’s free to use or modify.

A new open model rivals OpenAI’s o1, and it’s free to use or modify.

A new model from Hangzhou upstart DeepSeek delivers outstanding performance and may change the equation for training costs.

A new model from Hangzhou upstart DeepSeek delivers outstanding performance and may change the equation for training costs.

Merging multiple fine-tuned models is a less expensive alternative to hosting multiple specialized models. But, while model merging can deliver higher average performance across several tasks, it often results in lower performance on specific tasks. New work addresses this issue.

Merging multiple fine-tuned models is a less expensive alternative to hosting multiple specialized models. But, while model merging can deliver higher average performance across several tasks, it often results in lower performance on specific tasks. New work addresses this issue.

Harvard University amassed a huge new text corpus for training machine learning models.

Harvard University amassed a huge new text corpus for training machine learning models.

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

Anthropic analyzed 1 million anonymized conversations between users and Claude 3.5 Sonnet. The study found that most people used the model for software development and also revealed malfunctions and jailbreaks.

Anthropic analyzed 1 million anonymized conversations between users and Claude 3.5 Sonnet. The study found that most people used the model for software development and also revealed malfunctions and jailbreaks.

In 2025, I expect progress in training foundation models to slow down as we hit scaling limits and inference costs continue to rise.

In 2025, I expect progress in training foundation models to slow down as we hit scaling limits and inference costs continue to rise.
Stay updated with weekly AI News and Insights delivered to your inbox