Mistral’s Vision-Language Contender: Mistral unveils Pixtral Large, a rival to top vision-language models
Large Language Models (LLMs)

Mistral’s Vision-Language Contender: Mistral unveils Pixtral Large, a rival to top vision-language models

Mistral AI unveiled Pixtral Large, which rivals top models at processing combinations of text and images.

December 4, 20242 min read
Agents Open the Wallet: Stripe builds ecommerce agent toolkit for AI to securely spend money
Large Language Models (LLMs)

Agents Open the Wallet: Stripe builds ecommerce agent toolkit for AI to securely spend money

One of the world’s biggest payment processors is enabling large language models to spend real money.

December 4, 20241 min read
Agents Open the Wallet: Stripe builds ecommerce agent toolkit for AI to securely spend money
Large Language Models (LLMs)

Agents Open the Wallet: Stripe builds ecommerce agent toolkit for AI to securely spend money

One of the world’s biggest payment processors is enabling large language models to spend real money.

December 4, 20241 min read
AI Power Couple Recommits: Amazon deepens Anthropic partnership with $4 billion investment
Large Language Models (LLMs)

AI Power Couple Recommits: Amazon deepens Anthropic partnership with $4 billion investment

Amazon and Anthropic expanded their partnership, potentially strengthening Amazon Web Services’ AI infrastructure and lengthening the high-flying startup’s runway.

November 27, 20242 min read
AI Power Couple Recommits: Amazon deepens Anthropic partnership with $4 billion investment
Large Language Models (LLMs)

AI Power Couple Recommits: Amazon deepens Anthropic partnership with $4 billion investment

Amazon and Anthropic expanded their partnership, potentially strengthening Amazon Web Services’ AI infrastructure and lengthening the high-flying startup’s runway.

November 27, 20242 min read
Reasoning Revealed: DeepSeek-R1, a transparent challenger to OpenAI o1
Large Language Models (LLMs)

Reasoning Revealed: DeepSeek-R1, a transparent challenger to OpenAI o1

An up-and-coming Hangzhou AI lab unveiled a model that implements run-time reasoning similar to OpenAI o1 and delivers competitive performance. Unlike o1, it displays its reasoning steps.

November 27, 20242 min read
Reasoning Revealed: DeepSeek-R1, a transparent challenger to OpenAI o1
Large Language Models (LLMs)

Reasoning Revealed: DeepSeek-R1, a transparent challenger to OpenAI o1

An up-and-coming Hangzhou AI lab unveiled a model that implements run-time reasoning similar to OpenAI o1 and delivers competitive performance. Unlike o1, it displays its reasoning steps.

November 27, 20242 min read
More-Efficient Training for Transformers: Researchers reduce transformer training costs by 20% with minimal performance loss
Large Language Models (LLMs)

More-Efficient Training for Transformers: Researchers reduce transformer training costs by 20% with minimal performance loss

Researchers cut the processing required to train transformers by around 20 percent with only a slight degradation in performance.

November 20, 20242 min read
More-Efficient Training for Transformers: Researchers reduce transformer training costs by 20% with minimal performance loss
Large Language Models (LLMs)

More-Efficient Training for Transformers: Researchers reduce transformer training costs by 20% with minimal performance loss

Researchers cut the processing required to train transformers by around 20 percent with only a slight degradation in performance.

November 20, 20242 min read
Next-Gen Models Show Limited Gains: AI giants rethink model training strategy as scaling laws break down
Large Language Models (LLMs)

Next-Gen Models Show Limited Gains: AI giants rethink model training strategy as scaling laws break down

Builders of large AI models have relied on the idea that bigger neural networks trained on more data and given more processing power would show steady improvements. Recent developments are challenging that idea.

November 20, 20244 min read
Next-Gen Models Show Limited Gains: AI giants rethink model training strategy as scaling laws break down
Large Language Models (LLMs)

Next-Gen Models Show Limited Gains: AI giants rethink model training strategy as scaling laws break down

Builders of large AI models have relied on the idea that bigger neural networks trained on more data and given more processing power would show steady improvements. Recent developments are challenging that idea.

November 20, 20244 min read
Claude Controls Computers: Anthropic empowers Claude Sonnet 3.5 to operate desktop apps, but cautions remain
Large Language Models (LLMs)

Claude Controls Computers: Anthropic empowers Claude Sonnet 3.5 to operate desktop apps, but cautions remain

API commands for Claude Sonnet 3.5 enable Anthropic’s large language model to operate desktop apps much like humans do. Be cautious, though: It’s a work in progress.

November 6, 20243 min read
Claude Controls Computers: Anthropic empowers Claude Sonnet 3.5 to operate desktop apps, but cautions remain
Large Language Models (LLMs)

Claude Controls Computers: Anthropic empowers Claude Sonnet 3.5 to operate desktop apps, but cautions remain

API commands for Claude Sonnet 3.5 enable Anthropic’s large language model to operate desktop apps much like humans do. Be cautious, though: It’s a work in progress.

November 6, 20243 min read
Mistral AI Sharpens the Edge: Mistral AI unveils Ministral 3B and 8B models, outperforming rivals in small-scale AI
Large Language Models (LLMs)

Mistral AI Sharpens the Edge: Mistral AI unveils Ministral 3B and 8B models, outperforming rivals in small-scale AI

Mistral AI launched two models that raise the bar for language models with 8 billion or fewer parameters, small enough to run on many edge devices.

October 23, 20242 min read
Mistral AI Sharpens the Edge: Mistral AI unveils Ministral 3B and 8B models, outperforming rivals in small-scale AI
Large Language Models (LLMs)

Mistral AI Sharpens the Edge: Mistral AI unveils Ministral 3B and 8B models, outperforming rivals in small-scale AI

Mistral AI launched two models that raise the bar for language models with 8 billion or fewer parameters, small enough to run on many edge devices.

October 23, 20242 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox