The Proof Is in the Network: A transformer model that generates mathematical proofs
Transformer

The Proof Is in the Network: A transformer model that generates mathematical proofs

OpenAI’s Generative Pre-Trained Transformer (GPT) architecture has created coherent essays, images, and code. Now it generates mathematical proofs as well.

November 11, 20202 min read
GPT-3 Is No MD: GPT-3 lacks medical problem solving skills.
Transformer

GPT-3 Is No MD: GPT-3 lacks medical problem solving skills.

The world’s most sophisticated language model won’t replace your doctor anytime soon. Researchers at Nabla, an AI-enabled healthcare platform, found that GPT-3 lacks the logical reasoning skills to be a useful medical chatbot.

November 4, 20201 min read
GPT-3 Is No MD: GPT-3 lacks medical problem solving skills.
Transformer

GPT-3 Is No MD: GPT-3 lacks medical problem solving skills.

The world’s most sophisticated language model won’t replace your doctor anytime soon. Researchers at Nabla, an AI-enabled healthcare platform, found that GPT-3 lacks the logical reasoning skills to be a useful medical chatbot.

November 4, 20201 min read
The AI Community Splinters: Could geopolitics drive a wedge in the AI community?
Transformer

The AI Community Splinters: Could geopolitics drive a wedge in the AI community?

Will international rivalries fragment international cooperation in machine learning? Countries competing for AI dominance will lash out at competitors.

October 28, 20202 min read
The AI Community Splinters: Could geopolitics drive a wedge in the AI community?
Transformer

The AI Community Splinters: Could geopolitics drive a wedge in the AI community?

Will international rivalries fragment international cooperation in machine learning? Countries competing for AI dominance will lash out at competitors.

October 28, 20202 min read
More Efficient Transformers: BigBird is an efficient attention mechanism for transformers.
Transformer

More Efficient Transformers: BigBird is an efficient attention mechanism for transformers.

As transformer networks move to the fore in applications from language to vision, the time it takes them to crunch longer sequences becomes a more pressing issue. A new method lightens the computational load using sparse attention.

September 23, 20202 min read
More Efficient Transformers: BigBird is an efficient attention mechanism for transformers.
Transformer

More Efficient Transformers: BigBird is an efficient attention mechanism for transformers.

As transformer networks move to the fore in applications from language to vision, the time it takes them to crunch longer sequences becomes a more pressing issue. A new method lightens the computational load using sparse attention.

September 23, 20202 min read
Do Muppets Have Common Sense?: The Bert NLP model scores high on common sense test.
Transformer

Do Muppets Have Common Sense?: The Bert NLP model scores high on common sense test.

Two years after it pointed a new direction for language models, Bert still hovers near the top of several natural language processing leaderboards. A new study considers whether Bert simply excels at tracking word order or or learns something closer to common sense.

September 16, 20202 min read
Toward 1 Trillion Parameters: Microsoft upgrades its DeepSpeed optimization library.
Transformer

Toward 1 Trillion Parameters: Microsoft upgrades its DeepSpeed optimization library.

An open source library could spawn trillion-parameter neural networks and help small-time developers build big-league models. Microsoft upgraded DeepSpeed, a library that accelerates the PyTorch deep learning framework.

September 16, 20202 min read
Toward 1 Trillion Parameters: Microsoft upgrades its DeepSpeed optimization library.
Transformer

Toward 1 Trillion Parameters: Microsoft upgrades its DeepSpeed optimization library.

An open source library could spawn trillion-parameter neural networks and help small-time developers build big-league models. Microsoft upgraded DeepSpeed, a library that accelerates the PyTorch deep learning framework.

September 16, 20202 min read
Do Muppets Have Common Sense?: The Bert NLP model scores high on common sense test.
Transformer

Do Muppets Have Common Sense?: The Bert NLP model scores high on common sense test.

Two years after it pointed a new direction for language models, Bert still hovers near the top of several natural language processing leaderboards. A new study considers whether Bert simply excels at tracking word order or or learns something closer to common sense.

September 16, 20202 min read
The Transformation Continues: Technique boosts transformer performance on long sequences.
Transformer

The Transformation Continues: Technique boosts transformer performance on long sequences.

Transformer networks are gaining popularity as a high-accuracy alternative to recurrent neural networks. But they can run slowly when they’re applied to long sequences.

September 2, 20202 min read
The Transformation Continues: Technique boosts transformer performance on long sequences.
Transformer

The Transformation Continues: Technique boosts transformer performance on long sequences.

Transformer networks are gaining popularity as a high-accuracy alternative to recurrent neural networks. But they can run slowly when they’re applied to long sequences.

September 2, 20202 min read
Transforming Pixels: An image generation model using the GPT architecture
Transformer

Transforming Pixels: An image generation model using the GPT architecture

Language models like Bert, Ernie, and Elmo have achieved spectacular results based on clever pre-training approaches. New research applies some of those Sesame Street lessons into image processing.

August 5, 20202 min read
Transforming Pixels: An image generation model using the GPT architecture
Transformer

Transforming Pixels: An image generation model using the GPT architecture

Language models like Bert, Ernie, and Elmo have achieved spectacular results based on clever pre-training approaches. New research applies some of those Sesame Street lessons into image processing.

August 5, 20202 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox