Multimodal

40 Posts

Google’s Multimodal Challenger: All you need to know about Gemini, Google's new multimodal model
Multimodal

Google’s Multimodal Challenger: All you need to know about Gemini, Google's new multimodal model

Google unveiled Gemini, its bid to catch up to, and perhaps surpass, OpenAI’s GPT-4. Google demonstrated the Gemini family of models that accept any combination of text (including code), images, video, and audio and output text and images. The demonstrations and metrics were impressive...

December 13, 20233 min read
Google’s Multimodal Challenger: All you need to know about Gemini, Google's new multimodal model
Multimodal

Google’s Multimodal Challenger: All you need to know about Gemini, Google's new multimodal model

Google unveiled Gemini, its bid to catch up to, and perhaps surpass, OpenAI’s GPT-4. Google demonstrated the Gemini family of models that accept any combination of text (including code), images, video, and audio and output text and images. The demonstrations and metrics were impressive...

December 13, 20233 min read
Synthetic Data Helps Image Generators: OpenAI researchers improved text-to-image prompt following with generated captions.
Multimodal

Synthetic Data Helps Image Generators: OpenAI researchers improved text-to-image prompt following with generated captions.

Text-to-image generators often miss details in text prompts, and sometimes they misunderstand parts of a prompt entirely. Synthetic captions can help them follow prompts more closely.

November 1, 20233 min read
Synthetic Data Helps Image Generators: OpenAI researchers improved text-to-image prompt following with generated captions.
Multimodal

Synthetic Data Helps Image Generators: OpenAI researchers improved text-to-image prompt following with generated captions.

Text-to-image generators often miss details in text prompts, and sometimes they misunderstand parts of a prompt entirely. Synthetic captions can help them follow prompts more closely.

November 1, 20233 min read
Vision and Language Tightly Bound: Training on a single loss function improves multimiodal AI.
Multimodal

Vision and Language Tightly Bound: Training on a single loss function improves multimiodal AI.

Recent multimodal models process both text and images as sequences of tokens, but they learn to represent these distinct data types using separate loss functions. Recent work unifies the loss function as well.

March 15, 20232 min read
Vision and Language Tightly Bound: Training on a single loss function improves multimiodal AI.
Multimodal

Vision and Language Tightly Bound: Training on a single loss function improves multimiodal AI.

Recent multimodal models process both text and images as sequences of tokens, but they learn to represent these distinct data types using separate loss functions. Recent work unifies the loss function as well.

March 15, 20232 min read
GPT-4 Has Landed: Everything you need to know about GPT-4.
Multimodal

GPT-4 Has Landed: Everything you need to know about GPT-4.

Get ready for the next wave of language-model mania. OpenAI introduced the latest in its GPT series of large language models to widespread excitement. The company showed statistics and examples designed to demonstrate...

March 15, 20233 min read
GPT-4 Has Landed: Everything you need to know about GPT-4.
Multimodal

GPT-4 Has Landed: Everything you need to know about GPT-4.

Get ready for the next wave of language-model mania. OpenAI introduced the latest in its GPT series of large language models to widespread excitement. The company showed statistics and examples designed to demonstrate...

March 15, 20233 min read
One Model Does It All: Multi-task AI models got more sophisticated in 2022.
Multimodal

One Model Does It All: Multi-task AI models got more sophisticated in 2022.

Individual deep learning models proved their mettle in hundreds of tasks. The scope of multi-task models expanded dramatically in the past year.

December 21, 20222 min read
One Model Does It All: Multi-task AI models got more sophisticated in 2022.
Multimodal

One Model Does It All: Multi-task AI models got more sophisticated in 2022.

Individual deep learning models proved their mettle in hundreds of tasks. The scope of multi-task models expanded dramatically in the past year.

December 21, 20222 min read
One Model, Hundreds of Tasks: Multimodal Transformer Performs Over 600 Different Tasks
Multimodal

One Model, Hundreds of Tasks: Multimodal Transformer Performs Over 600 Different Tasks

Researchers took a step toward achieving a longstanding goal: One model that performs a whole lot of very different tasks. Scott Reed, Konrad Żołna, Emilio Parisotto and a team at DeepMind announced Gato.

May 18, 20222 min read
One Model, Hundreds of Tasks: Multimodal Transformer Performs Over 600 Different Tasks
Multimodal

One Model, Hundreds of Tasks: Multimodal Transformer Performs Over 600 Different Tasks

Researchers took a step toward achieving a longstanding goal: One model that performs a whole lot of very different tasks. Scott Reed, Konrad Żołna, Emilio Parisotto and a team at DeepMind announced Gato.

May 18, 20222 min read
AI Versus the Garbage Heap: How Amazon uses AI to cut waste.
Multimodal

AI Versus the Garbage Heap: How Amazon uses AI to cut waste.

Amazon reported long-term success using machine learning to shrink its environmental footprint. The online retailer developed a system that fuses product descriptions, images, and structured data to decide how an item should be packed for shipping.

January 12, 20222 min read
AI Versus the Garbage Heap: How Amazon uses AI to cut waste.
Multimodal

AI Versus the Garbage Heap: How Amazon uses AI to cut waste.

Amazon reported long-term success using machine learning to shrink its environmental footprint. The online retailer developed a system that fuses product descriptions, images, and structured data to decide how an item should be packed for shipping.

January 12, 20222 min read
Multimodal AI Takes Off: Multimodal Models, such as CLIP and DALL·E, are taking over AI.
Multimodal

Multimodal AI Takes Off: Multimodal Models, such as CLIP and DALL·E, are taking over AI.

While models like GPT-3 and EfficientNet, which work on text and images respectively, are responsible for some of deep learning’s highest-profile successes, approaches that find relationships between text and images made impressive

December 22, 20211 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox