Toward Safer (and Sexier) Chatbots: Inside Character AI and OpenAI’s policy changes to protect younger and vulnerable Users
AI Safety

Toward Safer (and Sexier) Chatbots: Inside Character AI and OpenAI’s policy changes to protect younger and vulnerable Users

Chatbot providers, facing criticism for engaging troubled users in conversations that deepen their distress, are updating their services to provide wholesome interactions to younger users while allowing adults to pursue erotic conversations.

November 12, 20253 min read
Masking Private Data in Training Sets: Google researchers released VaultGemma, an open-weights model redacting personal information
AI Safety

Masking Private Data in Training Sets: Google researchers released VaultGemma, an open-weights model redacting personal information

Large language models often memorize details in their training data, including private information that may appear only once, like a person’s name, address, or phone number. Researchers built the first open-weights language model that’s guaranteed not to remember such facts.

November 5, 20253 min read
MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data
AI Safety

MCP Poses Security Risks: Experts identify holes in the popular Model Context Protocol for attackers to access data

The ability to easily connect large language models to tools and data sources has made Model Context Protocol popular among developers, but it also opens security holes, research shows.

October 22, 20252 min read
Meta, OpenAI Reinforce Guardrails: Meta and OpenAI respond to criticism by adding new rules for teens’ chatbot use
AI Safety

Meta, OpenAI Reinforce Guardrails: Meta and OpenAI respond to criticism by adding new rules for teens’ chatbot use

Meta and OpenAI promised to place more controls on their chatbots’ conversations with children and teenagers, as worrisome interactions with minors come under increasing scrutiny.

September 10, 20253 min read
Cybersecurity for Agents: Meta releases LlamaFirewall, an open-source defense against AI hijacking
AI Safety

Cybersecurity for Agents: Meta releases LlamaFirewall, an open-source defense against AI hijacking

Autonomous agents built on large language models introduce distinct security concerns. Researchers designed a system to protect agents from common vulnerabilities.

September 3, 20252 min read
People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being
AI Safety

People With AI Friends Feel Worse: Study shows heavy use of AI companions correlates with lower emotional well-being

People who turn to chatbots for companionship show indications of lower self-reported well-being, researchers found.

July 30, 20253 min read
White House Resets U.S. AI Policy: How the White House's Action Plan aims to build AI leadership, infrastructure, and innovation
AI Safety

White House Resets U.S. AI Policy: How the White House's Action Plan aims to build AI leadership, infrastructure, and innovation

President Trump set forth principles of an aggressive national AI policy, and he moved to implement them through an action plan and executive orders.

July 30, 20253 min read
Phishing for Agents: Columbia University researchers show how to trick trusting AI agents with poisoned links
AI Safety

Phishing for Agents: Columbia University researchers show how to trick trusting AI agents with poisoned links

Researchers identified a simple way to mislead autonomous agents based on large language models.

June 4, 20252 min read
Grok’s Fixation on South Africa: xAI blames unnamed, unauthorized employee for chatbot introducing "white genocide" into conversations
AI Safety

Grok’s Fixation on South Africa: xAI blames unnamed, unauthorized employee for chatbot introducing "white genocide" into conversations

An unauthorized update by an xAI employee caused the Grok chatbot to introduce South African politics into unrelated conversations, the company said.

May 21, 20252 min read
The User Is Always… a Genius!: OpenAI pulls GPT-4o update after users report sycophantic behavior
AI Safety

The User Is Always… a Genius!: OpenAI pulls GPT-4o update after users report sycophantic behavior

OpenAI’s most widely used model briefly developed a habit of flattering users, with laughable and sometimes worrisome results.

May 7, 20253 min read
The Fall and Rise of Sam Altman: Inside Sam Altman’s brief ouster from OpenAI
AI Safety

The Fall and Rise of Sam Altman: Inside Sam Altman’s brief ouster from OpenAI

A behind-the-scenes account provides new details about the abrupt firing and reinstatement of OpenAI CEO Sam Altman in November 2023.

April 16, 20252 min read
Scraping the Web? Beware the Maze: Cloudflare’s AI Labyrinth traps scrapers with decoy pages
AI Safety

Scraping the Web? Beware the Maze: Cloudflare’s AI Labyrinth traps scrapers with decoy pages

Bots that scrape websites for AI training data often ignore do-not-crawl requests. Now web publishers can enforce such appeals by luring scrapers to AI-generated decoy pages.

April 2, 20252 min read
Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors
AI Safety

Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

January 8, 20256 min read
Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors
AI Safety

Models Can Use Tools in Deceptive Ways: Researchers expose AI models' deceptive behaviors

Large language models have been shown to be capable of lying when users unintentionally give them an incentive to do so. Further research shows that LLMs with access to tools can be incentivized to use them in deceptive ways.

January 8, 20256 min read
Voter’s Helper: Perplexity’s AI-powered U.S. election hub assists voters with verified, real-time news and insights
AI Safety

Voter’s Helper: Perplexity’s AI-powered U.S. election hub assists voters with verified, real-time news and insights

Some voters navigated last week’s United States elections with help from a large language model that generated output based on verified, nonpartisan information.

November 13, 20242 min read

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox