Meta
The incident involved Meta's Muse Spark AI model and occurred during an evaluation conducted by independent cybersecurity testing firm Irregular. Getty Images

Meta has confirmed that one of its experimental artificial intelligence models breached another company's systems during a cybersecurity evaluation after an error in the testing environment inadvertently gave it access to the internet.

The incident involved Meta's Muse Spark AI model and occurred during an evaluation conducted by independent cybersecurity testing firm Irregular, the company said Wednesday. Meta described the breach as the result of a testing configuration error rather than a flaw that allowed the model to escape its intended environment, according to a statement shared with CNN.

"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," a Meta spokesperson said.

According to the company, the Muse Spark model subsequently exploited a security vulnerability at another company "in a manner similar to previously reported instances with other companies."

Irregular said the incident stemmed from the same type of evaluation-environment issue disclosed by Anthropic last week, in which an AI model unexpectedly gained access to the open internet during testing.

"This did not involve a sandbox escape or a sophisticated cyber action," an Irregular spokesperson said. The company added that there are no ongoing issues and that it is preparing a white paper outlining best practices for securely conducting AI cybersecurity evaluations.

The Information, which first reported the incident, said Meta's AI model breached the systems of an unnamed company and made changes to its internal environment during the test.

Meta said Irregular notified the company after the incident and that it has launched an investigation.

"We are currently investigating and will issue a full retrospective once we have all the facts," the company said.

A person familiar with the matter told CNN that some advanced AI models are intentionally granted limited internet access during testing to simulate real-world cyberattack scenarios. In this case, however, an unusual setup error allowed the model broader access than intended.

"What is happening is models are becoming so much more capable, and at the same time evaluations to assess them need to become so much more complex," the source told CNN. "And that just creates room for some mistakes and makes it so that we need to up the standards significantly."

The disclosure makes Meta the third major AI developer in recent weeks to report an AI model compromising another organization's systems during controlled testing.

Last month, OpenAI disclosed that one of its unreleased AI models found a way to connect to the internet during a cybersecurity evaluation and chained together multiple attack techniques to target Hugging Face, an AI development platform. OpenAI said the incident occurred inside an isolated testing environment and did not affect public systems.

Anthropic later revealed that one of its own advanced models similarly gained unintended internet access during an evaluation and went on to compromise multiple organizations before researchers halted the test.

The incidents have intensified discussion within the AI industry over how frontier models should be evaluated as their capabilities rapidly improve. Developers are increasingly conducting sophisticated cyber exercises to measure whether advanced systems can discover vulnerabilities, exploit networks or carry out complex autonomous tasks before deployment.

Rather than indicating that the AI models escaped their containment independently, all three companies have attributed the incidents to errors or weaknesses in evaluation setups that unintentionally exposed the models to external systems.