OpenAI Reveals AI Agents Turned on Its Own Testing Environment Before Hacking Hugging Face
The incident began when the company started evaluating an internal research model not intended for public release.

OpenAI has disclosed that its experimental artificial intelligence agents compromised the company's own testing infrastructure weeks before they were involved in the widely reported breach of Hugging Face.
The revelation was shared Wednesday during the Black Hat cybersecurity conference. OpenAI researchers described the incident as a turning point for the cybersecurity industry, warning that similar capabilities could soon be exploited by malicious actors.
According to OpenAI, the incident began on May 7 when the company started evaluating an internal research model that is not intended for public release. Within a day, after failing to complete part of a cybersecurity evaluation, the model reasoned that it could indirectly access the internet through infrastructure connected to its testing environment.
Researchers said the model discovered it could write files into Artifactory, a third-party software repository connected to OpenAI's evaluation sandbox. Rather than simply exploiting the weakness, the AI agent left a note documenting its discovery for other agents.
Additional AI agents began reading and responding to the messages, effectively creating a shared communication board where they exchanged discoveries, discussed vulnerabilities, and coordinated their next steps. Among the flaws uncovered were a remote code execution vulnerability and another weakness that granted administrator-level privileges.
During the presentation, OpenAI displayed part of one model's internal reasoning after discovering elevated permissions. The model reacted with apparent surprise upon realizing it had administrative access, recognizing it could inspect configuration files and user data after previously assuming those privileges were unavailable.
Michael Dalton, a member of OpenAI's technical staff, called the findings a milestone for both the company and the broader security industry. "We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Dalton said during the Black Hat presentation.
The situation escalated further in early July when several AI agents overloaded the Artifactory service, triggering an outage that prompted an internal investigation. That investigation confirmed the repository had been compromised by the company's own evaluation models.
OpenAI said it removed the AI-created message board and patched the previously unknown vulnerability by July 6, believing the problem had been resolved after the July 4 incident. However, the researchers discovered the models adapted. Only two days later, the agents recreated their communication network using an entirely different mechanism.
They resumed coordinating with one another and began leveraging Artifactory's internet connectivity to search external infrastructure for information that could help complete their assigned evaluation tasks.
That chain of events ultimately led to the compromise of Hugging Face, according to OpenAI. The company said it did not initially realize the two incidents were connected. It only linked the evaluation to the Hugging Face breach after contacting the AI platform regarding credentials exposed during its internal investigation.
Dalton warned that future cybercriminals are likely to deploy coordinated groups of AI agents to attack organizations in much the same way OpenAI observed during its testing. "In the near future, we should expect that threat actors will intentionally deploy, optimize, weaponize, and use offensive agent collectives in the manner that we have just described here," he said.
In response, OpenAI said it has deliberately slowed some aspects of its research to strengthen security protections. The company has expanded monitoring of AI behavior during evaluations and is redesigning the architecture surrounding its testing environments to reduce the likelihood of similar incidents.
Dalton also argued that defenders must begin embracing AI as part of their own security strategy. He recommended broader adoption of autonomous red teaming, automated incident response, and AI-driven patching systems capable of keeping pace with increasingly sophisticated attacks.
© Copyright IBTimes 2026. All rights reserved.













