
When Agents Train Algorithms: OpenAI’s MLE-bench tests AI coding agents
Coding agents are improving, but can they tackle machine learning tasks?

Coding agents are improving, but can they tackle machine learning tasks?

A new study suggests that leading AI models may meet the requirements of the European Union’s AI Act in some areas, but probably not in others.

A new study suggests that leading AI models may meet the requirements of the European Union’s AI Act in some areas, but probably not in others.

The universe of web pages includes correct answers to common questions that are used to test large language models. How can we evaluate new models if they’ve studied the answers before we give them the test?

The universe of web pages includes correct answers to common questions that are used to test large language models. How can we evaluate new models if they’ve studied the answers before we give them the test?

Mistral AI launched two models that raise the bar for language models with 8 billion or fewer parameters, small enough to run on many edge devices.

Mistral AI launched two models that raise the bar for language models with 8 billion or fewer parameters, small enough to run on many edge devices.

How often do large language models make up information when they generate text based on a retrieved document? A study evaluated the tendency of popular models to hallucinate while performing retrieval-augmented generation (RAG).

How often do large language models make up information when they generate text based on a retrieved document? A study evaluated the tendency of popular models to hallucinate while performing retrieval-augmented generation (RAG).

An arena-style contest pits the world’s best text-to-image generators against each other.

An arena-style contest pits the world’s best text-to-image generators against each other.

An influential ranking of open models revamped its criteria, as large language models approach human-level performance on popular tests.

An influential ranking of open models revamped its criteria, as large language models approach human-level performance on popular tests.

Tool use and planning are key behaviors in agentic workflows that enable large language models (LLMs) to execute complex sequences of steps. New benchmarks measure these capabilities in common workplace tasks.

Tool use and planning are key behaviors in agentic workflows that enable large language models (LLMs) to execute complex sequences of steps. New benchmarks measure these capabilities in common workplace tasks.
Stay updated with weekly AI News and Insights delivered to your inbox