Transformers

42 Posts

Kimi K3 Reveals How A Giant Frontier AI Model Works: Moonshot's latest model outshines all but GPT-5.6 Sol and Claude 5 Fable
Transformers

Kimi K3 Reveals How A Giant Frontier AI Model Works: Moonshot's latest model outshines all but GPT-5.6 Sol and Claude 5 Fable

Moonshot’s latest model leapfrogged the month-old GLM-5.2 and a host of proprietary competitors to finish just behind GPT-5.6 Sol and Claude Fable 5 on many benchmarks.

July 24, 20264 min read
Biological Molecules as Language: ESMFold2 approaches AlphaFold 3 performance but with an open, Transformer-based architecture
Transformers

Biological Molecules as Language: ESMFold2 approaches AlphaFold 3 performance but with an open, Transformer-based architecture

Google’s AlphaFold models pioneered the task of finding the shapes of biologically active molecules, opening new pathways for drug development.

June 26, 20264 min read
Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud
Transformers

Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud

The third generation of Apple Foundation Models — fruit of Apple’s collaboration with Google — introduces a variation on the mixture-of-experts architecture that runs on local devices. 

June 26, 20263 min read
Built-In Conversational Interactivity: Thinking Machines reveals its first interaction model, a new type of multimodal AI
Transformers

Built-In Conversational Interactivity: Thinking Machines reveals its first interaction model, a new type of multimodal AI

Conversational models typically wait for a turn before they respond.

May 22, 20264 min read
How Liquids and Gases Behave: A dynamic fluids model appears to solve transformers’ pixellation problem
Transformers

How Liquids and Gases Behave: A dynamic fluids model appears to solve transformers’ pixellation problem

Simulating complex physical systems through traditional numerical methods is slow and expensive, and simulations based on machine learning are usually specialized for a specific type of system, such as water in a pipe or atmosphere surrounding a planet.

April 10, 20263 min read
Dark DNA Unveiled: Google’s AlphaGenome interprets DNA that regulates genetic expression
Transformers

Dark DNA Unveiled: Google’s AlphaGenome interprets DNA that regulates genetic expression

An open-weights model could help scientists compare the impact of genetic variations, identify mutations that cause diseases, and develop treatments.

April 10, 20263 min read
Learning Long Context at Inference: Test-Time Training End-to-End (TTT-E2E) retrains model weights to handle long inputs
Transformers

Learning Long Context at Inference: Test-Time Training End-to-End (TTT-E2E) retrains model weights to handle long inputs

Large language models typically become less accurate and slower when they process longer contexts, but researchers enabled an LLM to keep accuracy stable and inference time constant as its context grew.

April 3, 20264 min read
A Single Tokenizer for Visual Media: Apple’s AToken, a multimodal model with a single encoder and tokenizer for images, videos, and 3D objects
Transformers

A Single Tokenizer for Visual Media: Apple’s AToken, a multimodal model with a single encoder and tokenizer for images, videos, and 3D objects

Multimodal models typically use different tokenizers to embed different media types, and different encoders when training to generate media rather than classify it.

March 20, 20264 min read
Sleep Signals Predict Illness: SleepFM detects signs of neurological disorders years before symptoms manifest
Transformers

Sleep Signals Predict Illness: SleepFM detects signs of neurological disorders years before symptoms manifest

Difficulty sleeping often precedes heart disease, psychiatric disorders, and many other illnesses. Researchers used data gathered during sleep studies to detect such conditions.

February 20, 20262 min read
Faster Reasoning at the Edge: Liquid AI’s small reasoning model mixes attention with convolutional layers for efficiency
Transformers

Faster Reasoning at the Edge: Liquid AI’s small reasoning model mixes attention with convolutional layers for efficiency

Reasoning models in the 1 to 2 billion-parameter range typically require more than 1 gigabyte of RAM to run. Liquid AI released one that runs in less than 900 megabytes, and does it with exceptional speed and efficiency.

February 20, 20263 min read
Recipe for Smaller, Capable Models: Mistral uses cascade distillation on Mistral 3 to build Ministral family
Transformers

Recipe for Smaller, Capable Models: Mistral uses cascade distillation on Mistral 3 to build Ministral family

Mistral compressed Mistral Small 3.1 into much smaller versions, yielding a family of relatively small, open-weights, vision-language models that perform better by some measures than competing models of similar size. The method combines pruning and distillation.

February 6, 20263 min read
Refining Words in Pictures: Z.ai’s GLM-Image blends transformer and diffusion architectures for better text in images
Transformers

Refining Words in Pictures: Z.ai’s GLM-Image blends transformer and diffusion architectures for better text in images

Image generators often mangle text. An open-weights model outperforms open and proprietary competitors in text rendering.

January 30, 20263 min read
Training Cars to Reason: Nvidia’s Alpamayo-R1 is a robotics-style reasoning model for autonomous vehicles
Transformers

Training Cars to Reason: Nvidia’s Alpamayo-R1 is a robotics-style reasoning model for autonomous vehicles

Chain-of-thought reasoning can help autonomous vehicles decide what to do next.

January 23, 20262 min read
Baidu’s Multimodal Bids: Giant Ernie 5 natively generates multiple media; Ernie-4.5-VL-28B-A3B-Thinking tops Vision-Language metrics
Transformers

Baidu’s Multimodal Bids: Giant Ernie 5 natively generates multiple media; Ernie-4.5-VL-28B-A3B-Thinking tops Vision-Language metrics

Baidu debuted two models: a lightweight, open-weights, vision-language model and a giant, proprietary, multimodal model built to take on U.S. competitors.

December 3, 20253 min read
Top Agentic Results, Open Weights: Kimi K2 Thinking outperforms proprietary models with new techniques for agentic tool use
Transformers

Top Agentic Results, Open Weights: Kimi K2 Thinking outperforms proprietary models with new techniques for agentic tool use

The latest open-weights large language model from Moonshot AI challenges top proprietary LLMs at agentic tasks by executing hundreds of tool calls sequentially and pausing to think between each.

November 19, 20254 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox