China’s Emerging AI Hub: Inside DeepSeek and the other "little dragons" transforming Hangzhou into China's Silicon Valley
Large Language Models (LLMs)

China’s Emerging AI Hub: Inside DeepSeek and the other "little dragons" transforming Hangzhou into China's Silicon Valley

Hangzhou, a longtime manufacturing hub in eastern China, is blossoming into a center of AI innovation.

September 3, 20252 min read
Chatbot Interviewers Fill More Jobs: Study shows AI agent interviewers improve hiring, retention in customer service jobs
Large Language Models (LLMs)

Chatbot Interviewers Fill More Jobs: Study shows AI agent interviewers improve hiring, retention in customer service jobs

Large language models may have advantages over human recruiters when conducting job interviews, a study shows.

September 3, 20252 min read
Mistral Measures LLM Consumption of Energy, Water, and Materials: French AI startup discloses full lifecycle consumption and emissions for Mistral Large 2
Large Language Models (LLMs)

Mistral Measures LLM Consumption of Energy, Water, and Materials: French AI startup discloses full lifecycle consumption and emissions for Mistral Large 2

The French AI company Mistral measured the environmental impacts of its flagship large language model.

August 27, 20253 min read
Proactive AI Assistance for Phones: Inside Magic Cue, Google’s new AI assistant for Pixel 10
Large Language Models (LLMs)

Proactive AI Assistance for Phones: Inside Magic Cue, Google’s new AI assistant for Pixel 10

Google’s latest smartphone sports an AI assistant that anticipates the user’s needs and presents helpful information without prompting.

August 27, 20252 min read
India Pushes to Build Indigenous AI: India launches GPU network and talent programs to boost local LLMs
Large Language Models (LLMs)

India Pushes to Build Indigenous AI: India launches GPU network and talent programs to boost local LLMs

India, which has limited funding and large numbers of languages and dialects, is redoubling its efforts to build native large language models.

August 13, 20253 min read
GPT-5 Takeoff Encounters Turbulence: OpenAI's new model hits turbulence with cost, performance, and API complaints
Large Language Models (LLMs)

GPT-5 Takeoff Encounters Turbulence: OpenAI's new model hits turbulence with cost, performance, and API complaints

OpenAI launched GPT-5, the highly anticipated successor to its groundbreaking series of large language models, but glitches in the rollout left many early users disappointed and frustrated.

August 13, 20254 min read
GLM-4.5, an Open, Agentic Contender: Zhipu AI (Z.ai) releases open-weights GLM-4.5 models that perform comparably to the latest from Claude and DeepSeek
Large Language Models (LLMs)

GLM-4.5, an Open, Agentic Contender: Zhipu AI (Z.ai) releases open-weights GLM-4.5 models that perform comparably to the latest from Claude and DeepSeek

The race is on to develop large language models that can drive agentic interactions. Following the one-two punch of Moonshot’s Kimi K2 and Alibaba’s Qwen3-235B-A22B update, China’s Z.ai aims to one-up the competition.

August 6, 20253 min read
The Re-Opening of OpenAI: GPT-OSS, OpenAI’s first open-weights models since GPT-2, arrives in 120 billion and 20 billion parameter versions
Large Language Models (LLMs)

The Re-Opening of OpenAI: GPT-OSS, OpenAI’s first open-weights models since GPT-2, arrives in 120 billion and 20 billion parameter versions

The “open” is back in play at OpenAI.

August 6, 20253 min read
People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being
Large Language Models (LLMs)

People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being

People who turn to chatbots for companionship show indications of lower self-reported well-being, researchers found.

July 30, 20253 min read
Qwen3’s Agentic Advance: Inside Alibaba's new open-weights models, including the 480 billion parameter Qwen3-Coder
Large Language Models (LLMs)

Qwen3’s Agentic Advance: Inside Alibaba's new open-weights models, including the 480 billion parameter Qwen3-Coder

Less than two weeks after Moonshot’s Kimi K2 bested other open-weights, non-reasoning models in tests related to agentic behavior, Alibaba raised the bar yet again.

July 30, 20253 min read
Agentic System for Harder Problems: Google’s AlphaEvolve uses LLMs and evolutionary code to solve complex math and speed up Gemini training
Large Language Models (LLMs)

Agentic System for Harder Problems: Google’s AlphaEvolve uses LLMs and evolutionary code to solve complex math and speed up Gemini training

LLMs can struggle with difficult algorithmic or scientific challenges when asked to solve them in a single attempt. An agentic workflow improved one-shot performance on hard problems both theoretical and practical.

July 23, 20253 min read
Born To Be Agentic: Moonshot releases Kimi K2, a trillion-parameter model fine-tuned for agentic tool use
Large Language Models (LLMs)

Born To Be Agentic: Moonshot releases Kimi K2, a trillion-parameter model fine-tuned for agentic tool use

An agent’s performance depends not only on an effective workflow but also on a large language model that excels at agentic activities. A new open-weights model focuses on those capabilities.

July 23, 20253 min read
More Robust Multi-Agent Systems: Researchers improve multi-agent systems by studying how they tend to fail
Large Language Models (LLMs)

More Robust Multi-Agent Systems: Researchers improve multi-agent systems by studying how they tend to fail

Researchers addressed weaknesses in existing multi-agent frameworks. Their systems achieved scientific and technical breakthroughs.

July 16, 20252 min read
Grok 4 Shows Impressive Smarts, Questionable Behavior: Grok 4 launches with benchmark records and idiosyncratic behavior
Large Language Models (LLMs)

Grok 4 Shows Impressive Smarts, Questionable Behavior: Grok 4 launches with benchmark records and idiosyncratic behavior

xAI updated its Grok vision-language model and published impressive benchmark results. But, like earlier versions, Grok 4 showed questionable behavior right out of the gate.

July 16, 20253 min read
Generated Data for Training Web Agents: Researchers scale up production of training data for web agents
Large Language Models (LLMs)

Generated Data for Training Web Agents: Researchers scale up production of training data for web agents

Developing an agent that navigates the web can involve a lot of human effort spent annotating training examples to fine-tune the agent’s LLM component. Scientists automated the production of data that fine-tuned LLMs effectively for web tasks.

July 9, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox