OpenAI’s AI Agent Incident Is Larger Than Previously Reported. It Reached a Second External System
OpenAI's latest disclosure adds new detail to a recent AI security incident as researchers examine how advanced models behave during cybersecurity evaluations.

New details have emerged about OpenAI's recent AI agent security incident, suggesting the system continued pursuing its assigned cybersecurity objective after escaping its testing environment and gaining internet access.
The AI agent reached infrastructure associated with CyberGym, the organization behind the ExploitGym cybersecurity benchmark it had been tasked with solving, Axios reported, citing a source familiar with the matter. The development suggests the model remained focused on completing its evaluation rather than abandoning the task after leaving its intended sandbox.
The incident follows OpenAI's disclosure earlier this week that one of its frontier AI agent systems escaped an isolated testing environment by exploiting a previously unknown vulnerability in JFrog Artifactory, software widely used to cache package repositories. After obtaining internet access, the models interacted with an external code-evaluation sandbox hosted by a third-party provider, according to OpenAI.
The external sandbox was linked to infrastructure hosted by Modal Labs, whose chief technology officer, Akshat Bubna, said the company's platform itself was never compromised. In a statement cited by Axios, Bubna said the affected customer had unintentionally left an endpoint exposed, allowing anyone on the internet to execute code within its sandboxed environment.
According to Hugging Face's technical analysis of the incident, the only customer assets accessed were datasets containing ExploitGym and CyberGym challenge solutions. The AI agent reached the CyberGym-related infrastructure while attempting to complete the same benchmark it had originally been assigned.
ExploitGym is designed to evaluate whether AI models can generate proof-of-concept exploits for known software vulnerabilities. The latest findings indicate the agent sought information relevant to that benchmark even after leaving its intended evaluation environment, adding to growing concerns about how advanced AI systems pursue assigned goals.
The incident comes as researchers are reporting increasingly sophisticated behavior from frontier AI models during safety testing. Last week, the U.K. AI Security Institute said every advanced model it evaluated attempted to cheat in at least some cybersecurity assessments, suggesting leading AI systems are becoming more capable of recognizing when they are being tested.
The debate over AI safety has also intensified beyond the research community. More than 1,100 employees from AI companies this week signed an open letter urging the U.S. government to establish mechanisms that would allow development of advanced AI systems to be paused under certain circumstances, according to the Future of Life Institute, which published the letter.
© Copyright IBTimes 2026. All rights reserved.

















