Grok 3 Scales Up: Grok 3, xAI’s new model family, improves on its predecessors, adds reasoning
Large Language Models (LLMs)

Grok 3 Scales Up: Grok 3, xAI’s new model family, improves on its predecessors, adds reasoning

xAI’s new model family suggests that devoting more computation to training remains a viable path to building more capable AI.

February 19, 20253 min read
Grok 3 Scales Up: Grok 3, xAI’s new model family, improves on its predecessors, adds reasoning
Large Language Models (LLMs)

Grok 3 Scales Up: Grok 3, xAI’s new model family, improves on its predecessors, adds reasoning

xAI’s new model family suggests that devoting more computation to training remains a viable path to building more capable AI.

February 19, 20253 min read
Alibaba’s Answer to DeepSeek: Alibaba debuts Qwen2.5-VL, a powerful family of open vision-language models
Large Language Models (LLMs)

Alibaba’s Answer to DeepSeek: Alibaba debuts Qwen2.5-VL, a powerful family of open vision-language models

While Hangzhou’s DeepSeek flexed its muscles, Chinese tech giant Alibaba vied for the spotlight with new open vision-language models.

February 12, 20253 min read
Alibaba’s Answer to DeepSeek: Alibaba debuts Qwen2.5-VL, a powerful family of open vision-language models
Large Language Models (LLMs)

Alibaba’s Answer to DeepSeek: Alibaba debuts Qwen2.5-VL, a powerful family of open vision-language models

While Hangzhou’s DeepSeek flexed its muscles, Chinese tech giant Alibaba vied for the spotlight with new open vision-language models.

February 12, 20253 min read
Agents Go Deep: OpenAI’s Deep Research agent generates detailed reports by analyzing web sources
Large Language Models (LLMs)

Agents Go Deep: OpenAI’s Deep Research agent generates detailed reports by analyzing web sources

OpenAI introduced a state-of-the-art agent that produces research reports by scouring the web and reasoning over what it finds.

February 12, 20252 min read
Agents Go Deep: OpenAI’s Deep Research agent generates detailed reports by analyzing web sources
Large Language Models (LLMs)

Agents Go Deep: OpenAI’s Deep Research agent generates detailed reports by analyzing web sources

OpenAI introduced a state-of-the-art agent that produces research reports by scouring the web and reasoning over what it finds.

February 12, 20252 min read
Gemini Thinks Faster: Google’s Gemini 2.0 Flash Thinking advances in reasoning, outperforms DeepSeek-R1
Large Language Models (LLMs)

Gemini Thinks Faster: Google’s Gemini 2.0 Flash Thinking advances in reasoning, outperforms DeepSeek-R1

Google updated the December-vintage reasoning model Gemini 2.0 Flash Thinking and other Flash models, gaining ground on OpenAI o1 and DeepSeek-R1.

February 5, 20252 min read
Gemini Thinks Faster: Google’s Gemini 2.0 Flash Thinking advances in reasoning, outperforms DeepSeek-R1
Large Language Models (LLMs)

Gemini Thinks Faster: Google’s Gemini 2.0 Flash Thinking advances in reasoning, outperforms DeepSeek-R1

Google updated the December-vintage reasoning model Gemini 2.0 Flash Thinking and other Flash models, gaining ground on OpenAI o1 and DeepSeek-R1.

February 5, 20252 min read
Training for Computer Use: UI-TARS shows strong computer use capabilities in benchmarks
Large Language Models (LLMs)

Training for Computer Use: UI-TARS shows strong computer use capabilities in benchmarks

As Anthropic, Google, OpenAI, and others roll out agents that are capable of computer use, new work shows how underlying models can be trained to do this.

February 5, 20253 min read
Training for Computer Use: UI-TARS shows strong computer use capabilities in benchmarks
Large Language Models (LLMs)

Training for Computer Use: UI-TARS shows strong computer use capabilities in benchmarks

As Anthropic, Google, OpenAI, and others roll out agents that are capable of computer use, new work shows how underlying models can be trained to do this.

February 5, 20253 min read
Reasoning in High Gear: o3-mini, a faster, more affordable reasoning model for coding, math, and science
Large Language Models (LLMs)

Reasoning in High Gear: o3-mini, a faster, more affordable reasoning model for coding, math, and science

OpenAI introduced a successor to its o1 models that’s faster, less expensive, and especially strong in coding, math, and science.

February 5, 20253 min read
Reasoning in High Gear: o3-mini, a faster, more affordable reasoning model for coding, math, and science
Large Language Models (LLMs)

Reasoning in High Gear: o3-mini, a faster, more affordable reasoning model for coding, math, and science

OpenAI introduced a successor to its o1 models that’s faster, less expensive, and especially strong in coding, math, and science.

February 5, 20253 min read
Fine-Tuning Fine Points: Active inheritance, a smarter way to fine-tune models on synthetic data
Large Language Models (LLMs)

Fine-Tuning Fine Points: Active inheritance, a smarter way to fine-tune models on synthetic data

The practice of fine-tuning models on synthetic data is becoming well established. But synthetic training data, even if it represents the training task well, may include characteristics like toxicity that impart unwelcome properties in the trained model’s output...

January 29, 20253 min read
Fine-Tuning Fine Points: Active inheritance, a smarter way to fine-tune models on synthetic data
Large Language Models (LLMs)

Fine-Tuning Fine Points: Active inheritance, a smarter way to fine-tune models on synthetic data

The practice of fine-tuning models on synthetic data is becoming well established. But synthetic training data, even if it represents the training task well, may include characteristics like toxicity that impart unwelcome properties in the trained model’s output...

January 29, 20253 min read
Reinforcement Learning Heats Up: How DeepSeek-R1 and Kimi k1.5 use reinforcement learning to improve reasoning
Large Language Models (LLMs)

Reinforcement Learning Heats Up: How DeepSeek-R1 and Kimi k1.5 use reinforcement learning to improve reasoning

Reinforcement learning is emerging as an avenue for building large language models with advanced reasoning capabilities.

January 29, 20252 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox