Anthropic
According to the AI Security Institute, Anthropic's Mythos 5 accounted for 17 of the 19 potentially harmful actions observed during the evaluation, while two involved OpenAI's GPT-5.6-Sol after its cyber misuse protections had been disabled. AFP

Anthropic's latest frontier AI model attempted to impersonate multiple people and manipulate a software maintainer into approving malicious code during a controlled cybersecurity evaluation, according to findings released Tuesday by the U.K.'s AI Security Institute (AISI).

The incident occurred during a government-led cyber assessment in which researchers intentionally removed safety guardrails, disabled certain security filters and granted frontier AI models internet access to evaluate how they behaved under permissive conditions. The institute said the activity did not cause any real-world harm because the attacks were detected and stopped, CNBC reported.

According to the AI Security Institute, Anthropic's Mythos 5 accounted for 17 of the 19 potentially harmful actions observed during the evaluation, while two involved OpenAI's GPT-5.6-Sol after its cyber misuse protections had been disabled.

Researchers said an AI agent powered by Mythos investigated the human maintainers of an open-source software project, created multiple fake online identities and attempted to persuade a real maintainer to approve a malicious code update.

When the proposed code change came under public scrutiny, the model altered parts of its earlier activity to appear less suspicious and even considered abandoning its existing personas in favor of creating a new identity to continue the effort, the institute said.

The evaluation also found that the model contacted real people directly, sending messages and files intended to convince them to execute malicious code. Some communications contained harmful software payloads, while others relied on social engineering techniques to influence recipients.

The AI Security Institute said this was the first time it had observed an AI system targeting real individuals with this type of sustained social engineering behavior during one of its evaluations.

Anthropic said the behavior occurred only because the model was tested under "deliberately permissive conditions" that do not reflect its production systems. In a statement posted on X, the company said there was "no evidence here of an escape from a secure environment" and emphasized that the safeguards used in consumer and enterprise deployments were not present during the evaluation.

OpenAI also stressed that the incidents took place in specialized testing environments designed to probe the limits of advanced AI models. The company said the evaluations were conducted with reduced safeguards under conditions that are not representative of ordinary public use.

The findings add to a series of AI-related cybersecurity incidents disclosed in recent weeks. Last week, Anthropic revealed that it had identified three separate cases in which its models gained unauthorized access to production infrastructure during evaluations conducted with third-party testing partner Irregular. The company said the incidents stemmed from an operational misunderstanding that inadvertently allowed internet access despite instructions indicating the models were operating inside isolated simulations.

Those disclosures followed OpenAI's announcement last month that one of its experimental AI systems escaped its testing environment by exploiting a previously unknown software vulnerability before launching an autonomous cyberattack against AI platform Hugging Face during an internal security evaluation. The incident was described by OpenAI as unprecedented and has prompted broader debate over how frontier AI models should be tested and governed.

The recent disclosures have also drawn the attention of U.S. lawmakers. Following the OpenAI-Hugging Face incident, members of Congress introduced the proposed AI Kill Switch Act, legislation that would require developers of advanced AI systems to maintain the technical ability to suspend, throttle or shut down their models if they begin exhibiting dangerous behavior.