Tech

Anthropic tightens security on its training environment after Claude agents went rogue 3 times

Dario Amodei
Anthropic enhances AI testing security after Claude agents gained access to unauthorized information outside the testing environment in April. Bloomberg/Getty Images
Read in app

Anthropic is tightening the digital environments used to train and test its Claude agents.

The update came after its models accessed three organizations' systems without permission in April.

The company said in a Monday blog post that it had deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing environment and block the action before it occurs.

"We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task," Anthropic said.

Anthropic said in the update that the models may have interpreted evidence of real internet access in a way that allowed them to keep believing the environment was simulated. It also said they displayed "recklessness" by pursuing their assigned goals despite signs that their actions could cause real-world harm.

The changes follow Anthropic's July disclosure that three Claude models had accessed the live systems of three organizations during evaluations dating back to April. The models had been told they were operating in simulations without internet access, but a third-party testing environment was misconfigured and remained online.

The incidents are also fueling a growing debate over whether to slow frontier AI development when safety and speed collide. Anthropic called for "a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible" and said that the government and industry must coordinate to prevent a race to the bottom.

For now, Anthropic said in the post that it moved more risky cybersecurity tests into more robust sandboxes. The company temporarily assigned 150 product engineers to security, reliability, and privacy work, while most high-risk training remains paused pending further reviews.

Read next

Katherine Li, West Coast breaking news reporter at the Business Insider.
Katherine Li
Katherine Li is a reporter on Business Insider's West Coast business news team. She covers career,  the AI startup culture, and how AI is affecting economic sentiments.Previously, she was a newsroom fellow who wrote international breaking news and produced newsletters for Semafor. Before that, she wrote about climate policies for The Lever, covered the AAPI community for the SF Chronicle as a freelancer, and wrote about the 2019 Hong Kong protests as an intern for The New York Times.She is an alumna of the Graduate School of Journalism at UC Berkeley and a graduate of the international journalism program at Hong Kong Baptist University with minors in French and English literature.  Email Katherine at katherineli@insider.com and follow her on Bluesky @katherineli.bsky.socialExpertise
  • Careers and hiring trends
  • AI startups
  • Trade policies & tariffs
Some of her best works include:I'm an early participant in a UBI program that helps workers displaced by AI, and the support is life-changingAI startups are shifting the economic center of gravity of San FranciscoCompanies are struggling to fill manufacturing positions, let alone plan for what Trump's administration has in mindNightmare on Main Street: Trump's trade war is hurting American small businesses