OpenAI Forges Chains of Thought: OpenAI’s o1 models excel in reasoning, outperform GPT-4o in math and coding
Machine Learning Research

OpenAI Forges Chains of Thought: OpenAI’s o1 models excel in reasoning, outperform GPT-4o in math and coding

Preliminary versions of OpenAI’s new model family were trained explicitly to think step-by-step, yielding outstanding marks in math, science, and coding — but users can’t see their reasoning steps.

September 18, 20243 min read
OpenAI Forges Chains of Thought: OpenAI’s o1 models excel in reasoning, outperform GPT-4o in math and coding
Machine Learning Research

OpenAI Forges Chains of Thought: OpenAI’s o1 models excel in reasoning, outperform GPT-4o in math and coding

Preliminary versions of OpenAI’s new model family were trained explicitly to think step-by-step, yielding outstanding marks in math, science, and coding — but users can’t see their reasoning steps.

September 18, 20243 min read
2D-to-3D Goes Mainstream: AI systems from Stability AI and Shutterstock transform 2D images into 3D meshes in seconds
Machine Learning Research

2D-to-3D Goes Mainstream: AI systems from Stability AI and Shutterstock transform 2D images into 3D meshes in seconds

Traditionally, building 3D meshes for gaming, animation, product design, architecture, and the like has been labor-intensive. Now the ability to generate 3D meshes from a single image is widely available.

September 11, 20242 min read
2D-to-3D Goes Mainstream: AI systems from Stability AI and Shutterstock transform 2D images into 3D meshes in seconds
Machine Learning Research

2D-to-3D Goes Mainstream: AI systems from Stability AI and Shutterstock transform 2D images into 3D meshes in seconds

Traditionally, building 3D meshes for gaming, animation, product design, architecture, and the like has been labor-intensive. Now the ability to generate 3D meshes from a single image is widely available.

September 11, 20242 min read
Balancing Web Data Distributions: Automated method organizes large datasets for better model performance
Machine Learning Research

Balancing Web Data Distributions: Automated method organizes large datasets for better model performance

Datasets that were scraped from the web tend to be unbalanced, meaning examples of some classes (say, cats) are plentiful while examples of others (say, caterpillars) are scarce.

September 11, 20243 min read
Balancing Web Data Distributions: Automated method organizes large datasets for better model performance
Machine Learning Research

Balancing Web Data Distributions: Automated method organizes large datasets for better model performance

Datasets that were scraped from the web tend to be unbalanced, meaning examples of some classes (say, cats) are plentiful while examples of others (say, caterpillars) are scarce.

September 11, 20243 min read
Making LLMs Explainable: Google’s Gemma Scope probes how large language models think
Machine Learning Research

Making LLMs Explainable: Google’s Gemma Scope probes how large language models think

Researchers have probed the inner workings of individual layers of large language models. A new tool applies this approach to all layers.

September 4, 20243 min read
Making LLMs Explainable: Google’s Gemma Scope probes how large language models think
Machine Learning Research

Making LLMs Explainable: Google’s Gemma Scope probes how large language models think

Researchers have probed the inner workings of individual layers of large language models. A new tool applies this approach to all layers.

September 4, 20243 min read
Models Ranked for Hallucinations: Measuring language model hallucinations during information retrieval
Machine Learning Research

Models Ranked for Hallucinations: Measuring language model hallucinations during information retrieval

How often do large language models make up information when they generate text based on a retrieved document? A study evaluated the tendency of popular models to hallucinate while performing retrieval-augmented generation (RAG). 

September 4, 20242 min read
Models Ranked for Hallucinations: Measuring language model hallucinations during information retrieval
Machine Learning Research

Models Ranked for Hallucinations: Measuring language model hallucinations during information retrieval

How often do large language models make up information when they generate text based on a retrieved document? A study evaluated the tendency of popular models to hallucinate while performing retrieval-augmented generation (RAG). 

September 4, 20242 min read
Long Context Gets Up to Speed: AI21 Labs’ Jamba 1.5 outpaces transformers in long-text processing
Machine Learning Research

Long Context Gets Up to Speed: AI21 Labs’ Jamba 1.5 outpaces transformers in long-text processing

A new model generates tokens faster than current transformers, especially when processing long inputs.

September 4, 20243 min read
Long Context Gets Up to Speed: AI21 Labs’ Jamba 1.5 outpaces transformers in long-text processing
Machine Learning Research

Long Context Gets Up to Speed: AI21 Labs’ Jamba 1.5 outpaces transformers in long-text processing

A new model generates tokens faster than current transformers, especially when processing long inputs.

September 4, 20243 min read
Multimodal to the Max: 4M-21 multimodal model excels in handling diverse input and output types
Machine Learning Research

Multimodal to the Max: 4M-21 multimodal model excels in handling diverse input and output types

Researchers introduced a model that handles an unprecedented number of input and output types, including many related to performing computer vision tasks.

August 28, 20243 min read
Multimodal to the Max: 4M-21 multimodal model excels in handling diverse input and output types
Machine Learning Research

Multimodal to the Max: 4M-21 multimodal model excels in handling diverse input and output types

Researchers introduced a model that handles an unprecedented number of input and output types, including many related to performing computer vision tasks.

August 28, 20243 min read
Agentic Coding Strides Forward:  Genie coding assistant outperforms competitors on SWE-bench by over 30 percent
Machine Learning Research

Agentic Coding Strides Forward: Genie coding assistant outperforms competitors on SWE-bench by over 30 percent

An agentic coding assistant boosted the state of the art in an important benchmark by more than 30 percent.

August 28, 20242 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox