Large Multimodal Models (LMMs)

39 Posts

Kimi K3 Reveals How A Giant Frontier AI Model Works: Moonshot's latest model outshines all but GPT-5.6 Sol and Claude 5 Fable
Large Multimodal Models (LMMs)

Kimi K3 Reveals How A Giant Frontier AI Model Works: Moonshot's latest model outshines all but GPT-5.6 Sol and Claude 5 Fable

Moonshot’s latest model leapfrogged the month-old GLM-5.2 and a host of proprietary competitors to finish just behind GPT-5.6 Sol and Claude Fable 5 on many benchmarks.

July 24, 20264 min read
One Model Talks, Another One Thinks: GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) on the backend
Large Multimodal Models (LMMs)

One Model Talks, Another One Thinks: GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) on the backend

ChatGPT’s voice mode now listens and speaks at the same time, passing harder questions posed to the conversational model to a reasoning model in the background.

July 17, 20265 min read
Google Pairs Nano Banana Update With Video API: Gemini's image (Nano Banana 2 Lite) and video (Gemini Omni Flash) models, built to work together
Large Multimodal Models (LMMs)

Google Pairs Nano Banana Update With Video API: Gemini's image (Nano Banana 2 Lite) and video (Gemini Omni Flash) models, built to work together

Google built a low-cost, high-throughput pipeline for developers working with media, combining Gemini’s fastest image model yet with a similarly speedy multimodal model that can turn images into video with synchronized sound.

July 10, 20264 min read
Gemini 3.5 Flash Pairs Smarts With Speed: Google's updated Flash levels up, approaching top models but raising prices
Large Multimodal Models (LMMs)

Gemini 3.5 Flash Pairs Smarts With Speed: Google's updated Flash levels up, approaching top models but raising prices

Google’s faster model brings substantive gains at a substantially higher price, part of a rising trend in prices per token.

May 29, 20264 min read
Built-In Conversational Interactivity: Thinking Machines reveals its first interaction model, a new type of multimodal AI
Large Multimodal Models (LMMs)

Built-In Conversational Interactivity: Thinking Machines reveals its first interaction model, a new type of multimodal AI

Conversational models typically wait for a turn before they respond.

May 22, 20264 min read
ByteDance Bids for Video Leadership: ByteDance adds state-of-the-art Seedance 2.0 video to Capcut, while OpenAI retreats
Large Multimodal Models (LMMs)

ByteDance Bids for Video Leadership: ByteDance adds state-of-the-art Seedance 2.0 video to Capcut, while OpenAI retreats

As OpenAI prepares to shut down Sora, ByteDance made its own video generation model available to hundreds of millions of users.

May 8, 20264 min read
Kimi K2.6 Challenges Open-Weights Champs: Kimi K2.6 matches open Qwen3.6 Max andDeepSeek V4, falls just behind top closed models.
Large Multimodal Models (LMMs)

Kimi K2.6 Challenges Open-Weights Champs: Kimi K2.6 matches open Qwen3.6 Max andDeepSeek V4, falls just behind top closed models.

Moonshot AI’s updated Kimi model handles longer autonomous coding sessions and scales up its multi-agent orchestration relative to its predecessor.

May 1, 20264 min read
GPT-5.5 Outperforms, Hallucinates: OpenAI’s latest model tops leaderboards for coding, visual puzzles, and overall intelligence
Large Multimodal Models (LMMs)

GPT-5.5 Outperforms, Hallucinates: OpenAI’s latest model tops leaderboards for coding, visual puzzles, and overall intelligence

The latest update of OpenAI’s flagship model sets new states of the art in important benchmarks but has difficulty distinguishing between what it does and doesn't know.

May 1, 20264 min read
Life After Llama: With Muse Spark, Meta pivots away from its open-weights Llama strategy
Large Multimodal Models (LMMs)

Life After Llama: With Muse Spark, Meta pivots away from its open-weights Llama strategy

Meta pivoted from its open-weights strategy to deliver a closed alternative.

April 17, 20263 min read
A Single Tokenizer for Visual Media: Apple’s AToken, a multimodal model with a single encoder and tokenizer for images, videos, and 3D objects
Large Multimodal Models (LMMs)

A Single Tokenizer for Visual Media: Apple’s AToken, a multimodal model with a single encoder and tokenizer for images, videos, and 3D objects

Multimodal models typically use different tokenizers to embed different media types, and different encoders when training to generate media rather than classify it.

March 20, 20264 min read
DeepSeek Snubs Nvidia for Huawei: DeepSeek made its upcoming 4.0 model available for performance testing to Chinese chipmakers but not U.S, ones
Large Multimodal Models (LMMs)

DeepSeek Snubs Nvidia for Huawei: DeepSeek made its upcoming 4.0 model available for performance testing to Chinese chipmakers but not U.S, ones

DeepSeek, the Chinese developer of outstanding open-weights models, has withheld an upcoming update of its flagship model from U.S. chip makers, a move that intensifies the AI rivalry between the U.S. and China.

March 20, 20262 min read
Qwen3.5 Outperforms Bigger Models, Leads Vision Benchmarks: Alibaba’s latest flagship models are open-weights MoE performers in sizes from less than 1B parameters
Large Multimodal Models (LMMs)

Qwen3.5 Outperforms Bigger Models, Leads Vision Benchmarks: Alibaba’s latest flagship models are open-weights MoE performers in sizes from less than 1B parameters

The Qwen3.5 family of open-weights vision-language models includes impressive larger models as well as a smaller one that outperforms an OpenAI open-weights model 10 times its size.

March 20, 20263 min read
GPT-5.4’s Higher Performance, Higher Price: OpenAI’s GPT-5.4 Pro and GPT-5.4 Thinking challenge Google’s Gemini 3.1 Pro Preview as best all-around AI model
Large Multimodal Models (LMMs)

GPT-5.4’s Higher Performance, Higher Price: OpenAI’s GPT-5.4 Pro and GPT-5.4 Thinking challenge Google’s Gemini 3.1 Pro Preview as best all-around AI model

OpenAI updated its flagship models, extending the ability to use tools and setting the state of the art on a handful of benchmarks, and priced them at the top of the market. Its coding and agentic abilities have enabled Codex, OpenAI’s competitor to Anthropic’s Claude Code, to leap ahead.

March 16, 20263 min read
Nano Banana 2 Ups Performance/Price: Gemini 3.1 Flash Image makes photo generation and edits easier and faster
Large Multimodal Models (LMMs)

Nano Banana 2 Ups Performance/Price: Gemini 3.1 Flash Image makes photo generation and edits easier and faster

Google launched a cheaper, faster successor to its flagship image generator, delivering greater interactivity at roughly half the price.

March 6, 20263 min read
Gemini Takes the Lead: Google releases Gemini 3.1 Pro in preview, tops Intelligence Index at same price
Large Multimodal Models (LMMs)

Gemini Takes the Lead: Google releases Gemini 3.1 Pro in preview, tops Intelligence Index at same price

Google updated its flagship Gemini model, topping several benchmarks while undercutting competitors on performance per dollar.

February 27, 20263 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox