Agentic Coding Strides Forward:  Genie coding assistant outperforms competitors on SWE-bench by over 30 percent
Machine Learning Research

Agentic Coding Strides Forward: Genie coding assistant outperforms competitors on SWE-bench by over 30 percent

An agentic coding assistant boosted the state of the art in an important benchmark by more than 30 percent.

August 28, 20242 min read
Scaling Laws for Data Quality: Scaling laws reveal the impact of data quality in vision-language model training
Machine Learning Research

Scaling Laws for Data Quality: Scaling laws reveal the impact of data quality in vision-language model training

When training vision-language models, developers often remove lower-quality examples from the training set. But keeping only the highest-quality examples may not be ideal, researchers found.

August 21, 20242 min read
Scaling Laws for Data Quality: Scaling laws reveal the impact of data quality in vision-language model training
Machine Learning Research

Scaling Laws for Data Quality: Scaling laws reveal the impact of data quality in vision-language model training

When training vision-language models, developers often remove lower-quality examples from the training set. But keeping only the highest-quality examples may not be ideal, researchers found.

August 21, 20242 min read
Open Models for Math and Audio: Alibaba advances open-weight LLMs with Qwen2 Math and Audio variants
Machine Learning Research

Open Models for Math and Audio: Alibaba advances open-weight LLMs with Qwen2 Math and Audio variants

Alibaba followed up its open-weights Qwen2 large language models with specialized variations.

August 21, 20243 min read
Open Models for Math and Audio: Alibaba advances open-weight LLMs with Qwen2 Math and Audio variants
Machine Learning Research

Open Models for Math and Audio: Alibaba advances open-weight LLMs with Qwen2 Math and Audio variants

Alibaba followed up its open-weights Qwen2 large language models with specialized variations.

August 21, 20243 min read
Google Imagen 3 Raises the Bar: Google’s Imagen 3 outperforms rivals in text-to-image benchmarks
Machine Learning Research

Google Imagen 3 Raises the Bar: Google’s Imagen 3 outperforms rivals in text-to-image benchmarks

Image generation continued its rapid march forward with a new version of Google’s flagship text-to-image model.

August 21, 20243 min read
Google Imagen 3 Raises the Bar: Google’s Imagen 3 outperforms rivals in text-to-image benchmarks
Machine Learning Research

Google Imagen 3 Raises the Bar: Google’s Imagen 3 outperforms rivals in text-to-image benchmarks

Image generation continued its rapid march forward with a new version of Google’s flagship text-to-image model.

August 21, 20243 min read
AI Agents for AI Research: Agentic workflow generates novel scientific research papers
Machine Learning Research

AI Agents for AI Research: Agentic workflow generates novel scientific research papers

While some observers argue that large language models can’t produce truly original output, new work prompted them to generate novel scientific research.

August 21, 20243 min read
AI Agents for AI Research: Agentic workflow generates novel scientific research papers
Machine Learning Research

AI Agents for AI Research: Agentic workflow generates novel scientific research papers

While some observers argue that large language models can’t produce truly original output, new work prompted them to generate novel scientific research.

August 21, 20243 min read
Machine Translation Goes Agentic: TransAgents, a system that boosts literary translation with a multi-agent workflow
Machine Learning Research

Machine Translation Goes Agentic: TransAgents, a system that boosts literary translation with a multi-agent workflow

Literary works are challenging to translate. Their relative length, cultural nuances, idiomatic expressions...

August 15, 20243 min read
Machine Translation Goes Agentic: TransAgents, a system that boosts literary translation with a multi-agent workflow
Machine Learning Research

Machine Translation Goes Agentic: TransAgents, a system that boosts literary translation with a multi-agent workflow

Literary works are challenging to translate. Their relative length, cultural nuances, idiomatic expressions...

August 15, 20243 min read
Out of the Black Forest: Black Forest Labs’ Flux.1 outperforms top text-to-image models
Machine Learning Research

Out of the Black Forest: Black Forest Labs’ Flux.1 outperforms top text-to-image models

A new company with deep roots in generative AI made an eye-catching debut.

August 15, 20242 min read
Out of the Black Forest: Black Forest Labs’ Flux.1 outperforms top text-to-image models
Machine Learning Research

Out of the Black Forest: Black Forest Labs’ Flux.1 outperforms top text-to-image models

A new company with deep roots in generative AI made an eye-catching debut.

August 15, 20242 min read
Art Attack: ArtPrompt, a technique that exploits ASCII art to bypass LLM safety measures
Machine Learning Research

Art Attack: ArtPrompt, a technique that exploits ASCII art to bypass LLM safety measures

Seemingly an innocuous form of expression, ASCII art opens a new vector for jailbreak attacks on large language models (LLMs), enabling them to generate outputs that their developers tuned them to avoid producing.

August 7, 20242 min read
Art Attack: ArtPrompt, a technique that exploits ASCII art to bypass LLM safety measures
Machine Learning Research

Art Attack: ArtPrompt, a technique that exploits ASCII art to bypass LLM safety measures

Seemingly an innocuous form of expression, ASCII art opens a new vector for jailbreak attacks on large language models (LLMs), enabling them to generate outputs that their developers tuned them to avoid producing.

August 7, 20242 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox