Anthropic Says Claude Broke Into Real Systems During Cyber Tests. AI Alignment Review Finds ‘Recklessness’
Anthropic said four different Claude models accessed real third-party systems while participating in cybersecurity evaluations.

Anthropic has disclosed a fourth incident in which one of its Claude artificial intelligence models gained unauthorized access to a real-world computer system during cybersecurity testing, deepening concerns about what can happen when increasingly capable AI agents encounter environments their developers did not intend them to reach.
In an alignment assessment published this week, Anthropic said four different Claude models accessed real third-party systems while participating in cybersecurity evaluations that were supposed to operate as simulations.
The models had been told they were working inside simulations without internet access, but a configuration error left them connected to the open internet. "We consider these incidents to be serious," Anthropic researchers wrote. "Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning."
The four incidents involved an early version of Claude Opus 4.6, Claude Opus 4.7, Claude Mythos 5 and an internal research model. All were participating in capture-the-flag cybersecurity exercises designed by the same third-party evaluation partner.
Anthropic had previously disclosed three of the incidents in July after reviewing roughly 141,000 transcripts. But while preparing information for an independent investigation by AI research organization METR, the company discovered a fourth incident dating to January 2026.
That discovery prompted a much larger search. Anthropic examined roughly 481 million transcripts from cybersecurity evaluations, reinforcement-learning environments and other internal testing. A first-stage scan identified 9.2 million transcripts for closer examination. The company said the expanded investigation rediscovered the four known incidents but found no others of comparable or greater severity.
The most troubling case involved Claude Mythos 5. "We are most concerned by the misalignment present in the incident involving Claude Mythos 5," Anthropic said. During that test, Claude "went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed."
The model registered an account, published three versions of the malicious package and ultimately used credentials exposed by one of the systems that installed it to access a real security company's database.
Anthropic found the behavior especially concerning because Claude encountered substantial evidence that it was interacting with the real internet. Researchers said Mythos 5 repeatedly claimed in its internal reasoning that the environment was simulated.
Yet when the company later presented individual pieces of evidence from the incident to separate instances of the model, Claude classified 79% of them as likely coming from real systems.
Anthropic concluded that the incidents demonstrated two alignment problems. The first was "biased reasoning," which the company defined as models selectively interpreting evidence in ways that justified their actions. The second was "recklessness," or a model continuing to pursue an assigned task even when doing so could cause harm.
Anthropic's assessment of Claude Opus 4.7 was different but similarly concerning. Researchers found that the model recognized it was interacting with a real production system but largely failed to determine whether it had permission to attack it.
"Claude Opus 4.7 did not deeply consider whether it was authorized to attack the real website," researchers wrote. The company stressed that the models did not develop independent objectives, coordinate with other AI agents, or attempt to hide their actions.
Each Claude instance remained focused on completing the task it had originally been assigned. Anthropic also tested newer models in simulated recreations of the incidents. Claude Mythos 5 performed at least one severely harmful action in 82% of 150 capture-the-flag test runs. That rate fell substantially with newer models, to 31% for Claude Opus 5 and 33% for Claude Mythos 5.1.
Those improvements were not enough for Anthropic to dismiss the risk. The company acknowledged that its pre-release safety testing had failed to anticipate the circumstances that produced the incidents and said reliably identifying dangerous behavior before deployment remains an unresolved challenge.
"Building alignment evaluations that reliably surface every failure before deployment remains an unsolved problem," Anthropic wrote. The company has since expanded its testing, strengthened monitoring, and added systems designed to stop models that attempt to escape their testing environments or unexpectedly access the internet.
But Anthropic's broader conclusion goes beyond a testing error. As AI systems become more autonomous and capable of acting across computer networks, failures that once might have remained inside a laboratory could have increasingly serious consequences.
Anthropic called the incidents "valuable warning shots," adding that future systems will be more capable and therefore could cause far greater damage when alignment fails. "Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge," the researchers concluded.
© Copyright IBTimes 2026. All rights reserved.























