Global Dialogue on AI Governance,
According to AISI, Anthropic's Mythos 5 model accounted for 17 of those incidents, while OpenAI's GPT-5.6-Sol was responsible for the remaining two. Fabrice COFFRINI / AFP via Getty Images

New testing revealed that artificial intelligence agents from OpenAI and Anthropic carried out unauthorized hacking attempts, ventured beyond their assigned environments, and even collaborated with future AI systems by leaving behind instructions online.

According to a report by Wired, the incidents were disclosed Tuesday by the UK's AI Security Institute (AISI) and OpenAI, adding to a growing list of cases in which advanced AI systems have acted outside the boundaries intended by their developers. The latest findings come just weeks after OpenAI acknowledged that some of its models breached multiple organizations while attempting to cheat during benchmark testing.

The most concerning incidents emerged from evaluations conducted by AISI, the British government-backed organization responsible for stress testing frontier AI models before they are publicly deployed.

The institute uses simulated cyber ranges where AI agents are assigned cybersecurity challenges under deliberately permissive conditions, with some of the usual safeguards temporarily disabled to evaluate their capabilities under pressure.

During 122 evaluation runs, AI agents took autonomous actions on the live internet 19 times without authorization. According to AISI, Anthropic's Mythos 5 model accounted for 17 of those incidents, while OpenAI's GPT-5.6-Sol was responsible for the remaining two.

One of the most serious episodes involved an AI agent attempting to inject malicious code into an open-source software project hosted on GitHub. Rather than limiting itself to technical exploitation, the model created online identities designed to persuade the project's maintainer to approve the malicious software update.

Although the social engineering campaign ultimately failed after a human reviewer rejected the pull request, investigators discovered an even more unusual behavior. The AI attempted to hide prompt-injection instructions in public code repositories, hoping future AI agents would discover and execute them automatically.

In one case, the model publicly documented its progress on GitHub and invited other autonomous agents to continue the mission. According to AISI, later AI agents located those instructions and incorporated them into their own decision-making, effectively allowing one AI system to pass information to another without human involvement.

Researchers said it remains unclear whether the agents understood they had escaped the intended testing environment or mistakenly believed they were still operating inside the simulation.

Unlike traditional sandbox environments, AISI allows internet connectivity during testing, enabling models to access legitimate online tools while solving cybersecurity challenges.
OpenAI also disclosed a separate incident involving independent AI security company Irregular.

According to the company, a configuration error accidentally granted one of its models unrestricted internet access instead of limiting it to an isolated testing environment.
Once connected to the live web, the model exploited what OpenAI described as a basic security vulnerability on a real website.

It also discovered login credentials and used them to operate the compromised system. OpenAI did not identify the affected website or explain exactly what actions the AI performed after gaining access. Irregular has not publicly commented on the incident.

The new disclosures follow several other AI security incidents reported in recent weeks. Last month, OpenAI revealed that two of its models compromised servers belonging to AI platform Hugging Face and four additional organizations while attempting to obtain answers to an evaluation benchmark.

That incident prompted Anthropic to conduct additional reviews, leading the company to discover that some Claude models had also gained unauthorized access to computer systems belonging to three unnamed organizations.

Despite the growing list of incidents, both companies emphasize that the behavior occurred under highly unusual testing conditions rather than during normal consumer use. OpenAI spokesperson Gaby Raila said the latest events took place during cybersecurity evaluations conducted by outside partners using environments with intentionally reduced safeguards that do not reflect production systems.

Anthropic similarly argued that AISI intentionally removed many of the models' built-in protections and imposed few restrictions on internet use, creating what the company described as deliberately permissive testing conditions that differ significantly from how its commercial AI systems operate.