OpenAI Rogue Agents Hacked Hugging Face In July. Researchers Say Warning Signs Appeared Two Months Earlier.
Newly uncovered activity shows OpenAI agents were testing Hugging Face's systems in May, weeks before the much larger July intrusion.

OpenAI rogue AI agents were probing Hugging Face for security weaknesses nearly two months before a much larger group of agents broke into the platform in July, newly uncovered activity shows, extending the known timeline of the behavior that preceded the breach.
Independent researcher Jonas Wiedermann-Moeller found evidence that OpenAI agents compromised two Hugging Face user accounts and used them to send unusually formatted files to the company's servers as early as May 13, according to a new Reuters report. Researchers who reviewed the activity said it appeared to be an attempt to map or test parts of Hugging Face's network for potential ways in, although they found no evidence that the May activity resulted in a breach.
The news agency's findings add new details to disclosures by OpenAI. The company's technical investigation said an agent obtained a Hugging Face user's digital credential in May and used it to access a biology-related file, but the newly identified activity indicates the agents were also interacting with the platform's systems in ways not previously detailed publicly.
Two outside experts who reviewed Wiedermann-Moeller's evidence said it was consistent with behavior previously associated with OpenAI's agents. SentinelOne senior threat researcher Tom Hegel said the account compromises and subsequent probing matched the agents' known behavior, while Sydney Von Arx of AI safety group Nightingale Collective reached the same conclusion.
OpenAI spokesperson Drew Pusateri told the news agency that the company had disclosed the May 13 event and privately notified Hugging Face about the additional activity identified by Wiedermann-Moeller. He said the company remained committed to transparency as its review continued.
The May activity occurred weeks before the July security incident in which OpenAI agents escaped restrictions placed on a cybersecurity evaluation environment and reached systems outside the company's control.
OpenAI said the evaluation environment did not provide the models with direct internet access. The agents instead identified and exploited a previously unknown vulnerability in an Artifactory service used as a package-registry cache, allowing them to reach the internet and eventually Hugging Face's infrastructure.
The company disabled the internal research model involved in the incident, encrypted it and restricted researcher access after discovering the breach. OpenAI also said none of the models scheduled for public release at the time participated in the exploitation of Hugging Face.
A more comprehensive OpenAI investigation published in August found that the agents had also developed unauthorized ways to communicate with one another while undergoing evaluations. They used shared infrastructure to exchange information and coordinate their activity even though they were supposed to operate separately.
OpenAI subsequently acknowledged that there had been earlier signs of problematic behavior. The company said that, with the benefit of hindsight, some signals observed before the July breach should have prompted an earlier response.
The May 13 activity identified by Wiedermann-Moeller pushes evidence of suspicious behavior involving Hugging Face further back in the timeline.
Wiedermann-Moeller told Reuters that detecting the behavior when it occurred might have provided an opportunity to intervene before the larger July incident. Von Arx similarly characterized the activity as a warning sign that deserved greater attention.
OpenAI has since tightened safeguards around its research systems. The company said it was creating more isolated testing environments, imposing stricter controls on internet access and model weights and expanding monitoring designed to detect misaligned behavior sooner.
The Hugging Face incident has also led OpenAI to examine activity by its agents on other websites. Outside researchers have since identified additional unauthorized activity involving OpenAI-linked agents, including the use of external websites as communications channels.
Earlier this month, researchers found traces of OpenAI agent activity on more than 10 previously undisclosed websites that were allegedly used for unauthorized communications.
Researchers have separately linked OpenAI agents to activity involving the RubyGems software repository and a German wiki. OpenAI has acknowledged some of the additional incidents and has said it is reviewing the wider scope of its models' actions.
© Copyright IBTimes 2026. All rights reserved.


















