Private Benchmarks for Fairer Tests: Scale AI launches SEAL leaderboards to benchmark model performance
Benchmarks

Private Benchmarks for Fairer Tests: Scale AI launches SEAL leaderboards to benchmark model performance

Scale AI offers new leaderboards based on its own benchmarks.

June 19, 20243 min read
Private Benchmarks for Fairer Tests: Scale AI launches SEAL leaderboards to benchmark model performance
Benchmarks

Private Benchmarks for Fairer Tests: Scale AI launches SEAL leaderboards to benchmark model performance

Scale AI offers new leaderboards based on its own benchmarks.

June 19, 20243 min read
Benchmarks for Industry: Vals AI evaluates large language models on industry-specific tasks.
Benchmarks

Benchmarks for Industry: Vals AI evaluates large language models on industry-specific tasks.

How well do large language models respond to professional-level queries in various industry domains? A new company aims to find out.

April 24, 20242 min read
Benchmarks for Industry: Vals AI evaluates large language models on industry-specific tasks.
Benchmarks

Benchmarks for Industry: Vals AI evaluates large language models on industry-specific tasks.

How well do large language models respond to professional-level queries in various industry domains? A new company aims to find out.

April 24, 20242 min read
Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots
Benchmarks

Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots

Training an agent that controls a robot arm to perform a task — say, opening a door — that involves a sequence of motions (reach, grasp, turn, pull, release) can take from tens of thousands to millions of examples...

July 12, 20233 min read
Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots
Benchmarks

Sample-Efficient Training for Robots: Reinforcement learning from human feedback to train robots

Training an agent that controls a robot arm to perform a task — say, opening a door — that involves a sequence of motions (reach, grasp, turn, pull, release) can take from tens of thousands to millions of examples...

July 12, 20233 min read
When Trees Outdo Neural Networks: Decision Trees Perform Best on Most Tabular Data
Benchmarks

When Trees Outdo Neural Networks: Decision Trees Perform Best on Most Tabular Data

While neural networks perform well on image, text, and audio datasets, they fall behind decision trees and their variations for tabular datasets. New research looked into why.

November 23, 20222 min read
When Trees Outdo Neural Networks: Decision Trees Perform Best on Most Tabular Data
Benchmarks

When Trees Outdo Neural Networks: Decision Trees Perform Best on Most Tabular Data

While neural networks perform well on image, text, and audio datasets, they fall behind decision trees and their variations for tabular datasets. New research looked into why.

November 23, 20222 min read
Humanized Training for Robot Arms: New Research Improves Robot Performance and Adaptability
Benchmarks

Humanized Training for Robot Arms: New Research Improves Robot Performance and Adaptability

Robots trained via reinforcement learning usually study videos of robots performing the task at hand. A new approach used videos of humans to pre-train robotic arms.

July 29, 20222 min read
Humanized Training for Robot Arms: New Research Improves Robot Performance and Adaptability
Benchmarks

Humanized Training for Robot Arms: New Research Improves Robot Performance and Adaptability

Robots trained via reinforcement learning usually study videos of robots performing the task at hand. A new approach used videos of humans to pre-train robotic arms.

July 29, 20222 min read
Toward Next-Gen Language Models: New Benchmarks Test the Limits of Large Language Models
Benchmarks

Toward Next-Gen Language Models: New Benchmarks Test the Limits of Large Language Models

A new benchmark aims to raise the bar for large language models. Researchers at 132 institutions worldwide introduced the Beyond the Imitation Game benchmark (BIG-bench), which includes tasks that humans perform well but current state-of-the-art models don’t.

June 22, 20222 min read
Toward Next-Gen Language Models: New Benchmarks Test the Limits of Large Language Models
Benchmarks

Toward Next-Gen Language Models: New Benchmarks Test the Limits of Large Language Models

A new benchmark aims to raise the bar for large language models. Researchers at 132 institutions worldwide introduced the Beyond the Imitation Game benchmark (BIG-bench), which includes tasks that humans perform well but current state-of-the-art models don’t.

June 22, 20222 min read
AI Progress Report: Stanford University's fifth annual AI Report for 2022
Benchmarks

AI Progress Report: Stanford University's fifth annual AI Report for 2022

A new study showcases AI’s growing importance worldwide. What’s new: The fifth annual AI Index from Stanford University’s Institute for Human-Centered AI documents rises in funding, regulation, and performance.

March 23, 20222 min read
AI Progress Report: Stanford University's fifth annual AI Report for 2022
Benchmarks

AI Progress Report: Stanford University's fifth annual AI Report for 2022

A new study showcases AI’s growing importance worldwide. What’s new: The fifth annual AI Index from Stanford University’s Institute for Human-Centered AI documents rises in funding, regulation, and performance.

March 23, 20222 min read
Transformer Variants Head to Head: A benchmark for comparing different AI transformers.
Benchmarks

Transformer Variants Head to Head: A benchmark for comparing different AI transformers.

The transformer architecture has inspired a plethora of variations. Yet researchers have used a patchwork of metrics to evaluate their performance, making them hard to compare. New work aims to level the playing field.

March 3, 20212 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox