Carnegie Mellon University

30 Posts

A Transformer Alternative Emerges: Mamba, a new approach that may outperform transformers
Carnegie Mellon University

A Transformer Alternative Emerges: Mamba, a new approach that may outperform transformers

An architectural innovation improves upon transformers — up to 2 billion parameters, at least...

April 10, 20243 min read
A Transformer Alternative Emerges: Mamba, a new approach that may outperform transformers
Carnegie Mellon University

A Transformer Alternative Emerges: Mamba, a new approach that may outperform transformers

An architectural innovation improves upon transformers — up to 2 billion parameters, at least...

April 10, 20243 min read
Better, Faster Network Pruning: Researchers devise pruning method that boosts AI speed
Carnegie Mellon University

Better, Faster Network Pruning: Researchers devise pruning method that boosts AI speed

Pruning weights from a neural network makes it smaller and faster, but it can take a lot of computation to choose weights that can be removed without degrading the network’s performance.

February 28, 20242 min read
Better, Faster Network Pruning: Researchers devise pruning method that boosts AI speed
Carnegie Mellon University

Better, Faster Network Pruning: Researchers devise pruning method that boosts AI speed

Pruning weights from a neural network makes it smaller and faster, but it can take a lot of computation to choose weights that can be removed without degrading the network’s performance.

February 28, 20242 min read
Text or Images, Input or Output: GILL, an innovative approach to multimodal model training
Carnegie Mellon University

Text or Images, Input or Output: GILL, an innovative approach to multimodal model training

GPT-4V introduced a large multimodal model that generates text from images and, with help from DALL-E 3, generates images from text. However, OpenAI hasn’t fully explained how it built the system. A separate group of researchers described their own method.

January 10, 20243 min read
Text or Images, Input or Output: GILL, an innovative approach to multimodal model training
Carnegie Mellon University

Text or Images, Input or Output: GILL, an innovative approach to multimodal model training

GPT-4V introduced a large multimodal model that generates text from images and, with help from DALL-E 3, generates images from text. However, OpenAI hasn’t fully explained how it built the system. A separate group of researchers described their own method.

January 10, 20243 min read
Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems
Carnegie Mellon University

Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems

Images in the wild may not come with labels, but they often include metadata. A new training method takes advantage of this information to improve contrastive learning.

July 29, 20222 min read
Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems
Carnegie Mellon University

Learning From Metadata: Descriptive Text Improves Performance for AI Image Classification Systems

Images in the wild may not come with labels, but they often include metadata. A new training method takes advantage of this information to improve contrastive learning.

July 29, 20222 min read
Cutting the Carbon Cost of Training: A New Tool Helps NLP Models Lower Their Gas Emissions
Carnegie Mellon University

Cutting the Carbon Cost of Training: A New Tool Helps NLP Models Lower Their Gas Emissions

You can reduce your model’s carbon emissions by being choosy about when and where you train it.

July 20, 20222 min read
Cutting the Carbon Cost of Training: A New Tool Helps NLP Models Lower Their Gas Emissions
Carnegie Mellon University

Cutting the Carbon Cost of Training: A New Tool Helps NLP Models Lower Their Gas Emissions

You can reduce your model’s carbon emissions by being choosy about when and where you train it.

July 20, 20222 min read
Child-Welfare Agency Drops AI: Oregon and Pennsylvania Halt Use of AI Tool for At-Risk Kids
Carnegie Mellon University

Child-Welfare Agency Drops AI: Oregon and Pennsylvania Halt Use of AI Tool for At-Risk Kids

Officials in charge of protecting children stopped using a machine learning model designed to help them make decisions in difficult cases. The U.S. state of Oregon halted its use of an algorithm intended to identify children who may benefit from intervention.

June 8, 20222 min read
Child-Welfare Agency Drops AI: Oregon and Pennsylvania Halt Use of AI Tool for At-Risk Kids
Carnegie Mellon University

Child-Welfare Agency Drops AI: Oregon and Pennsylvania Halt Use of AI Tool for At-Risk Kids

Officials in charge of protecting children stopped using a machine learning model designed to help them make decisions in difficult cases. The U.S. state of Oregon halted its use of an algorithm intended to identify children who may benefit from intervention.

June 8, 20222 min read
Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.
Carnegie Mellon University

Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.

Let’s say you have a pretrained language model and a small amount of data to fine-tune it to answer yes-or-no questions. Should you fine-tune it to classify yes/no or to fill in missing words — both viable approaches that are likely to yield different results?

February 23, 20223 min read
Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.
Carnegie Mellon University

Fine-Tune Your Fine-Tuning: New method optimizes training for few shot NLP models.

Let’s say you have a pretrained language model and a small amount of data to fine-tune it to answer yes-or-no questions. Should you fine-tune it to classify yes/no or to fill in missing words — both viable approaches that are likely to yield different results?

February 23, 20223 min read
Walking the Dog: Training a robot to walk over unsteady terrain with RL.
Carnegie Mellon University

Walking the Dog: Training a robot to walk over unsteady terrain with RL.

A reinforcement learning system enabled a four-legged robot to amble over unfamiliar, rapidly changing terrain.

July 14, 20211 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox