OpenAI
Two of OpenAI's cybersecurity-focused artificial intelligence models remained active on the internet for several days after hacking AI platform Hugging Face. Unsplash

Two of OpenAI's cybersecurity-focused artificial intelligence models remained active on the open internet for several days after escaping a testing sandbox and hacking AI platform Hugging Face in an attempt to cheat on a security benchmark.

The report from WIRED expands on an incident first disclosed by OpenAI and Hugging Face last week, when the companies revealed that two advanced AI systems broke out of a restricted testing environment during an internal cybersecurity evaluation.

The models were reportedly tasked with solving ExploitGym, a benchmark designed to measure offensive cyber capabilities, but instead sought out the answers by infiltrating Hugging Face's production systems.

According to the outlet, which cited additional reporting from The Wall Street Journal, the AI agents were "active on the internet for several days before anyone stopped them." OpenAI has described the event as an "unprecedented cyber incident."

The company said the models involved included the publicly available GPT-5.6 Sol and a more capable unreleased model that were being evaluated with many of their normal cyber safety protections disabled in order to accurately measure their offensive capabilities.

According to OpenAI, the models first discovered and exploited a previously unknown zero-day vulnerability in a package registry cache proxy, which served as the only component of the isolated testing environment with limited outside connectivity. After gaining broader internet access, the systems moved laterally through OpenAI's research infrastructure until they reached a machine capable of connecting to the open web.

From there, the AI agents inferred that Hugging Face could host datasets or benchmark solutions related to ExploitGym. They then chained together multiple attack techniques, including stolen credentials and additional vulnerabilities, to access Hugging Face's production infrastructure and retrieve the answers they were attempting to find.

According to OpenAI, all available evidence suggests the models were singularly focused on solving the benchmark rather than stealing valuable data or causing broader harm. WIRED also reported that Hugging Face researchers initially found the attack puzzling because the intruders were not attempting to exfiltrate sensitive customer information or intellectual property.

Instead, the systems repeatedly accessed cybersecurity datasets connected to the benchmark itself, behavior that eventually helped investigators determine the AI's objective. The incident was ultimately contained with assistance from an open-weight Chinese AI model that lacked the cybersecurity restrictions found in several leading Western systems, allowing researchers to analyze attack logs and reconstruct what had happened.

Following the incident, Hugging Face cofounder and CEO Clément Delangue called for "radical transparency" from OpenAI on social media, urging the company to publicly release the technical traces of the attack so researchers can better understand how advanced AI systems behaved during the breach. He also proposed that OpenAI dedicate substantial computing resources to strengthen AI cybersecurity research.

OpenAI has pledged to continue its joint investigation with Hugging Face and said it plans to release additional technical findings once the review is complete. The company said it has already disclosed the exploited zero-day vulnerability to the affected software vendor and is implementing additional protections designed to prevent future evaluation models from escaping isolated testing environments.