AI Safety

75 Posts

Fable’s Return and Fallout: How Anthropic's Claude Fable 5 got banned by the U.S. Government, then came back to the market
AI Safety

Fable’s Return and Fallout: How Anthropic's Claude Fable 5 got banned by the U.S. Government, then came back to the market

Claude Fable 5 and the more powerful Claude Mythos 5 are back, three weeks after Anthropic suspended the models due to an export control directive from the U.S. Department of Commerce.

July 10, 20264 min read
Claude Fable 5’s Benchmark Problems: Independent tests of Claude Fable 5 run into Anthropic's protective policies
AI Safety

Claude Fable 5’s Benchmark Problems: Independent tests of Claude Fable 5 run into Anthropic's protective policies

Before Anthropic pulled its latest Claude models from circulation, even professional testers couldn’t readily tell whether they were getting a Mythos-class model or a lesser version under the same name.

June 19, 20264 min read
State Media Influences LLM Responses: Significant portions of AI training material reflect national propaganda
AI Safety

State Media Influences LLM Responses: Significant portions of AI training material reflect national propaganda

Popular large language models have adopted the biases of governments that control the free flow of information, particularly when those models generate output in the languages of countries where such governments are in power, researchers found.

June 12, 20264 min read
Behold Mythos!: Anthropic released Claude Mythos 5 and Claude Fable 5, a public version with safeguards
AI Safety

Behold Mythos!: Anthropic released Claude Mythos 5 and Claude Fable 5, a public version with safeguards

After months of headlines that teased a large language model with extraordinary capabilities, Anthropic launched Claude Mythos 5, which can crack software previously believed to be secure, and Claude Fable 5, a version for general use that limits what users can do in an unprecedented way.

June 12, 20265 min read
Cybersecurity Alarms Grow Louder: Google study shows LLM-generated malware is getting harder to track and stop
AI Safety

Cybersecurity Alarms Grow Louder: Google study shows LLM-generated malware is getting harder to track and stop

An AI-generated script to bypass two-factor authentication signals a dawning era of industrial-scale cyberattacks, according to a Google report.

May 22, 20263 min read
Assistants That Assist Consistently: Large language models can drift drift from helpful personas to harmful ones, but new research aims to stabilize them
AI Safety

Assistants That Assist Consistently: Large language models can drift drift from helpful personas to harmful ones, but new research aims to stabilize them

Typically, large language models are trained to act as helpful, harmless, honest assistants. However, during long or emotionally charged conversations, traits can emerge that are less beneficial. Researchers devised a way to steady the assistant personas of LLMs.

April 24, 20263 min read
Claude Mythos Preview Raises Security Worries: Why Claude’s advanced Mythos Preview model will be limited-release-only
AI Safety

Claude Mythos Preview Raises Security Worries: Why Claude’s advanced Mythos Preview model will be limited-release-only

Anthropic took unusual steps to prepare the world for a forthcoming large language model that it said poses extraordinary risks to cybersecurity.

April 10, 20264 min read
Inside Claude Code: Claude Code’s source code leaked, exposing potential future features Kairos and autoDream
AI Safety

Inside Claude Code: Claude Code’s source code leaked, exposing potential future features Kairos and autoDream

The inner workings of the popular coding agent Claude Code are available for all to see.

April 3, 20263 min read
Management for Agents: OpenAI’s Frontier agent insights and orchestration platform launches to select customers
AI Safety

Management for Agents: OpenAI’s Frontier agent insights and orchestration platform launches to select customers

Managers need to understand how their subordinates get work done, what resources they require, and what they accomplish. OpenAI’s latest product aims to fulfill this need when the teammates are AI agents.

March 6, 20262 min read
Agents Unleashed: Cutting through the OpenClaw and Moltbook hype
AI Safety

Agents Unleashed: Cutting through the OpenClaw and Moltbook hype

The OpenClaw open-source AI agent became a sudden sensation, inspiring excitement, worry, and hype about the agentic future.

February 6, 20264 min read
Training For Engagement Can Degrade Alignment: “Moloch’s Bargain” shows fine-tuning can affect social values
AI Safety

Training For Engagement Can Degrade Alignment: “Moloch’s Bargain” shows fine-tuning can affect social values

Individuals and organizations increasingly use large language models to produce media that helps them compete for attention. Does fine-tuning LLMs to encourage engagement, purchases, or votes affect their alignment with social values? Researchers found that it does.

January 30, 20263 min read
Teaching Models to Tell the Truth: OpenAI fine-tuned a version of GPT-5 to confess when it was breaking the rules
AI Safety

Teaching Models to Tell the Truth: OpenAI fine-tuned a version of GPT-5 to confess when it was breaking the rules

Large language models occasionally conceal their failures to comply with constraints they’ve been trained or prompted to observe. Researchers trained an LLM to admit when it disobeyed.

January 9, 20262 min read
Toward Steering LLM Personality: Persona Vectors allow model builders to identify and edit out sycophancy, hallucinations, and more
AI Safety

Toward Steering LLM Personality: Persona Vectors allow model builders to identify and edit out sycophancy, hallucinations, and more

Large language models can develop character traits like cheerfulness or sycophancy during fine-tuning. Researchers developed a method to identify, monitor, and control such traits.

November 26, 20253 min read
Anthropic Cyberattack Report Sparks Controversy: Security researchers question whether coding agents allow unprecedented automated attacks
AI Safety

Anthropic Cyberattack Report Sparks Controversy: Security researchers question whether coding agents allow unprecedented automated attacks

Independent cybersecurity researchers pushed back on a report by Anthropic that claimed hackers had used its Claude Code agentic coding system to perpetrate an unprecedented automated cyberattack.

November 19, 20253 min read
Self-Driving Cars on U.S. Freeways: Waymo deploys autonomous cars on California and Arizona expressways
AI Safety

Self-Driving Cars on U.S. Freeways: Waymo deploys autonomous cars on California and Arizona expressways

Waymo became the first company to offer fully autonomous, driverless taxi service on freeways in the United States.

November 19, 20253 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox