Attention for Image Generation: Combining GANs and transformers for more believable images.
Machine Learning Research

Attention for Image Generation: Combining GANs and transformers for more believable images.

Attention quantifies how each part of one input affects the various parts of another. Researchers added a step that reverses this comparison to produce more convincing images.

April 7, 20212 min read
Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.
Machine Learning Research

Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.

It’s well established that pretraining a model on a large dataset improves performance on fine-tuned tasks. In sufficient quantity and paired with a big model, even data scraped from the internet at random can contribute to the performance boost.

April 7, 20212 min read
Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.
Machine Learning Research

Pretraining on Uncurated Data: How unlabeled data improved computer vision accuracy.

It’s well established that pretraining a model on a large dataset improves performance on fine-tuned tasks. In sufficient quantity and paired with a big model, even data scraped from the internet at random can contribute to the performance boost.

April 7, 20212 min read
Pictures From Words and Gestures: AI model generates captions as users mouse over images.
Machine Learning Research

Pictures From Words and Gestures: AI model generates captions as users mouse over images.

A new system combines verbal descriptions and crude lines to visualize complex scenes. Google researchers led by Jing Yu Koh proposed Tag-Retrieve-Compose-Synthesize (TReCS), a system that generates photorealistic images by describing what they want to see while mousing around on a blank screen.

March 31, 20212 min read
Pictures From Words and Gestures: AI model generates captions as users mouse over images.
Machine Learning Research

Pictures From Words and Gestures: AI model generates captions as users mouse over images.

A new system combines verbal descriptions and crude lines to visualize complex scenes. Google researchers led by Jing Yu Koh proposed Tag-Retrieve-Compose-Synthesize (TReCS), a system that generates photorealistic images by describing what they want to see while mousing around on a blank screen.

March 31, 20212 min read
Same Patient, Different Views: Contrastive pretraining improves medical imaging AI.
Machine Learning Research

Same Patient, Different Views: Contrastive pretraining improves medical imaging AI.

When you lack labeled training data, pretraining a model on unlabeled data can compensate. New research pretrained a model three times to boost performance on a medical imaging task.

March 31, 20212 min read
Vision Models Get Some Attention: Researchers add self-attention to convolutional neural nets.
Machine Learning Research

Vision Models Get Some Attention: Researchers add self-attention to convolutional neural nets.

Self-attention is a key element in state-of-the-art language models, but it struggles to process images because its memory requirement rises rapidly with the size of the input. New research addresses the issue with a simple twist on a convolutional neural network.

March 31, 20212 min read
Same Patient, Different Views: Contrastive pretraining improves medical imaging AI.
Machine Learning Research

Same Patient, Different Views: Contrastive pretraining improves medical imaging AI.

When you lack labeled training data, pretraining a model on unlabeled data can compensate. New research pretrained a model three times to boost performance on a medical imaging task.

March 31, 20212 min read
Vision Models Get Some Attention: Researchers add self-attention to convolutional neural nets.
Machine Learning Research

Vision Models Get Some Attention: Researchers add self-attention to convolutional neural nets.

Self-attention is a key element in state-of-the-art language models, but it struggles to process images because its memory requirement rises rapidly with the size of the input. New research addresses the issue with a simple twist on a convolutional neural network.

March 31, 20212 min read
Deciphering The Brain’s Visual Signals: AI uses brain activity to create images.
Machine Learning Research

Deciphering The Brain’s Visual Signals: AI uses brain activity to create images.

What’s creepier than images from the sci-fi TV series Doctor Who? Images generated by a network designed to visualize what goes on in peoples’ brains while they watch Doctor Who.

March 24, 20212 min read
Deciphering The Brain’s Visual Signals: AI uses brain activity to create images.
Machine Learning Research

Deciphering The Brain’s Visual Signals: AI uses brain activity to create images.

What’s creepier than images from the sci-fi TV series Doctor Who? Images generated by a network designed to visualize what goes on in peoples’ brains while they watch Doctor Who.

March 24, 20212 min read
Good Labels for Cropped Images: AI technique adds text labels to random image crops.
Machine Learning Research

Good Labels for Cropped Images: AI technique adds text labels to random image crops.

In training an image recognition model, it’s not uncommon to augment the data by cropping original images randomly. But if an image contains several objects, a cropped version may no longer match its label. Researchers developed a way to make sure random crops are labeled properly.

March 17, 20212 min read
Good Labels for Cropped Images: AI technique adds text labels to random image crops.
Machine Learning Research

Good Labels for Cropped Images: AI technique adds text labels to random image crops.

In training an image recognition model, it’s not uncommon to augment the data by cropping original images randomly. But if an image contains several objects, a cropped version may no longer match its label. Researchers developed a way to make sure random crops are labeled properly.

March 17, 20212 min read
Seeing People From a New Angle: Neural Body is an AI tool for generating 3D images of people.
Machine Learning Research

Seeing People From a New Angle: Neural Body is an AI tool for generating 3D images of people.

The University of Hong Kong, and Cornell University to create Neural Body, a procedure that generates novel views of a single human character based on shots from only a few angles.

March 10, 20212 min read
Seeing People From a New Angle: Neural Body is an AI tool for generating 3D images of people.
Machine Learning Research

Seeing People From a New Angle: Neural Body is an AI tool for generating 3D images of people.

The University of Hong Kong, and Cornell University to create Neural Body, a procedure that generates novel views of a single human character based on shots from only a few angles.

March 10, 20212 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox