Better Image Processing Through Self-Supervised Learning: Meta’s DINOv3 gets an updated loss term and improved vision performance
Machine Learning Research

Better Image Processing Through Self-Supervised Learning: Meta’s DINOv3 gets an updated loss term and improved vision performance

DINOv2 showed that a vision transformer pretrained on unlabeled images could produce embeddings that are useful for a wide variety of tasks. Now it has been updated to improve the performance of its embeddings in segmentation and other vision tasks.

August 27, 20253 min read
Does Your Model Generalize or Memorize: Researchers find models with more parameters copy more bits from training sets
Machine Learning Research

Does Your Model Generalize or Memorize: Researchers find models with more parameters copy more bits from training sets

Benchmarks can measure how well large language models apply what they’ve learned from their training data to new data, but it’s harder to measure the degree to which they simply memorized their training data. New work proposes a way to gauge memorization.

August 20, 20253 min read
Mixture of Video Experts: Alibaba’s Wan 2.2 video models adopt a new architecture to sort noisy from less-noisy inputs
Machine Learning Research

Mixture of Video Experts: Alibaba’s Wan 2.2 video models adopt a new architecture to sort noisy from less-noisy inputs

The mixture-of-experts approach that has boosted the performance of large language models may do the same for video generation.

August 20, 20253 min read
Training Data for Coding Assistants: Stanford and Alibaba build bug fixing dataset and pipeline to train AI
Machine Learning Research

Training Data for Coding Assistants: Stanford and Alibaba build bug fixing dataset and pipeline to train AI

A bottleneck in fine-tuning large language models for software engineering is building a dataset that can show them how to edit code, search for subroutines, write test scripts, control a terminal, manage a file system, and so on. Researchers built a pipeline that produces such data automatically.

August 13, 20252 min read
GPT-5 Takeoff Encounters Turbulence: OpenAI's new model hits turbulence with cost, performance, and API complaints
Machine Learning Research

GPT-5 Takeoff Encounters Turbulence: OpenAI's new model hits turbulence with cost, performance, and API complaints

OpenAI launched GPT-5, the highly anticipated successor to its groundbreaking series of large language models, but glitches in the rollout left many early users disappointed and frustrated.

August 13, 20254 min read
Robot Surgeon Cuts and Clips: Doctors at Stanford, Johns Hopkins, and Optosurgical operate on animal organs without human intervention
Machine Learning Research

Robot Surgeon Cuts and Clips: Doctors at Stanford, Johns Hopkins, and Optosurgical operate on animal organs without human intervention

An autonomous robot performed intricate surgical operations without human intervention.

August 6, 20254 min read
GLM-4.5, an Open, Agentic Contender: Zhipu AI (Z.ai) releases open-weights GLM-4.5 models that perform comparably to the latest from Claude and DeepSeek
Machine Learning Research

GLM-4.5, an Open, Agentic Contender: Zhipu AI (Z.ai) releases open-weights GLM-4.5 models that perform comparably to the latest from Claude and DeepSeek

The race is on to develop large language models that can drive agentic interactions. Following the one-two punch of Moonshot’s Kimi K2 and Alibaba’s Qwen3-235B-A22B update, China’s Z.ai aims to one-up the competition.

August 6, 20253 min read
The Re-Opening of OpenAI: GPT-OSS, OpenAI’s first open-weights models since GPT-2, arrives in 120 billion and 20 billion parameter versions
Machine Learning Research

The Re-Opening of OpenAI: GPT-OSS, OpenAI’s first open-weights models since GPT-2, arrives in 120 billion and 20 billion parameter versions

The “open” is back in play at OpenAI.

August 6, 20253 min read
People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being
Machine Learning Research

People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being

People who turn to chatbots for companionship show indications of lower self-reported well-being, researchers found.

July 30, 20253 min read
Qwen3’s Agentic Advance: Inside Alibaba's new open-weights models, including the 480 billion parameter Qwen3-Coder
Machine Learning Research

Qwen3’s Agentic Advance: Inside Alibaba's new open-weights models, including the 480 billion parameter Qwen3-Coder

Less than two weeks after Moonshot’s Kimi K2 bested other open-weights, non-reasoning models in tests related to agentic behavior, Alibaba raised the bar yet again.

July 30, 20253 min read
Agentic System for Harder Problems: Google’s AlphaEvolve uses LLMs and evolutionary code to solve complex math and speed up Gemini training
Machine Learning Research

Agentic System for Harder Problems: Google’s AlphaEvolve uses LLMs and evolutionary code to solve complex math and speed up Gemini training

LLMs can struggle with difficult algorithmic or scientific challenges when asked to solve them in a single attempt. An agentic workflow improved one-shot performance on hard problems both theoretical and practical.

July 23, 20253 min read
Born To Be Agentic: Moonshot releases Kimi K2, a trillion-parameter model fine-tuned for agentic tool use
Machine Learning Research

Born To Be Agentic: Moonshot releases Kimi K2, a trillion-parameter model fine-tuned for agentic tool use

An agent’s performance depends not only on an effective workflow but also on a large language model that excels at agentic activities. A new open-weights model focuses on those capabilities.

July 23, 20253 min read
More Robust Multi-Agent Systems: Researchers improve multi-agent systems by studying how they tend to fail
Machine Learning Research

More Robust Multi-Agent Systems: Researchers improve multi-agent systems by studying how they tend to fail

Researchers addressed weaknesses in existing multi-agent frameworks. Their systems achieved scientific and technical breakthroughs.

July 16, 20252 min read
Grok 4 Shows Impressive Smarts, Questionable Behavior: Grok 4 launches with benchmark records and idiosyncratic behavior
Machine Learning Research

Grok 4 Shows Impressive Smarts, Questionable Behavior: Grok 4 launches with benchmark records and idiosyncratic behavior

xAI updated its Grok vision-language model and published impressive benchmark results. But, like earlier versions, Grok 4 showed questionable behavior right out of the gate.

July 16, 20253 min read
Generated Data for Training Web Agents: Researchers scale up production of training data for web agents
Machine Learning Research

Generated Data for Training Web Agents: Researchers scale up production of training data for web agents

Developing an agent that navigates the web can involve a lot of human effort spent annotating training examples to fine-tune the agent’s LLM component. Scientists automated the production of data that fine-tuned LLMs effectively for web tasks.

July 9, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox