Anthropic Spotted Unauthorized Actions By Agents. It Is Pausing Some Training And Evaluations.
OpenAI also paused work in August after concluding it could pose critical cybersecurity risks.

Anthropic said it paused work on some AI training and cybersecurity evaluations after spotting unauthorized actions by agents.
The company recalled in a blog post different incidents in which Claude models "gained unauthorized access to real computer systems" due to a "misconfiguration inside a third-party evaluation environment."
It also noted that the "UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet."
As a result, the company said, it is making changes and pausing external cyber evaluations of pre-released models because of the former incidents, claiming they "reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards)."
The company also paused higher-risk reinforcement learning environments on pre-relased models. Most of them have resumed, but some are paused pending manual review or newer monitoring tools.
"To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible," said the company, which added that is redirecting resources toward model security.
OpenAI made a similar decision in August after concluding a new model could pose critical cybersecurity risks.
In a social media publication, CEO Sam Altman said the decision will seek to "ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us."
"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," he added.
The suspension follows a high-profile incident in which the company disclosed a model had managed to break out of a sandbox environment and hack company Hugging Face in an attempt to achieve the testing goal.
"In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers," OpenAI disclosed.
After the incident and the fact that the Astra model potentially reached a critical threshold, the company said "the risks associated with developing and testing them internally also grow."
"Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling," OpenAI said, noting that this includes a "two-week pause in reinforcement learning (RL) training on our latest models intended for deployment."
© Copyright IBTimes 2026. All rights reserved.




















