AI
The Open Secure AI Alliance is developing the Shared AI Findings Exchange, a proposed incident-reporting framework designed to give the industry a common way to document rogue AI incidents. Getty Images

A coalition of more than 120 technology and cybersecurity organizations is pushing for a new system to track and disclose incidents involving rogue artificial intelligence agents, as autonomous AI tools are increasingly gaining the ability to interact with real computer systems with limited human supervision.

The Open Secure AI Alliance, whose members include Nvidia, Cisco and CrowdStrike, is developing the Shared AI Findings Exchange, or SAFE, a proposed incident-reporting framework designed to give the industry a common way to document AI failures, notify affected organizations and learn from cases in which agents cross security boundaries.

The effort comes amid mounting concern over "agentic AI," systems capable of carrying out multi-step tasks independent of human oversight. Under the draft SAFE rules, participating companies would be required to report incidents when an AI system accesses and/or modifies a third-party system without authorization.

Incidents would also qualify when an agent escapes a sandbox or other security boundary, gains access to confidential third-party information, or continues probing a production system after its operator knows or reasonably suspects the activity is unauthorized.

One of the proposal's most significant provisions makes clear that an AI developer's intentions would not determine whether an incident needs to be disclosed. "Intent does not determine whether an event is reportable," the draft says. "Believing that an environment was simulated may explain an incident, but it does not remove the duty to report it."

Members would also preserve detailed evidence from incidents and certain near misses, potentially including "prompts, agent traces, tool calls, logs, configurations, model and safeguard versions and third-party dependencies." The goal is to reconstruct exactly how an autonomous system behaved and identify safeguards that could prevent similar failures.

The reporting process would operate on a strict timeline. A directly affected organization would be notified as soon as possible, while customers facing credible exposure would be informed within 72 hours. Members would submit an initial confidential report to SAFE within four business days, with the framework also envisioning later public reporting and remediation updates when appropriate.

Justin Boitano, Nvidia's vice president and general manager of enterprise computing, compared the system to the aviation industry's approach to investigating accidents. "The way I think of it is the harness, which has visibility into everything the agent is doing, is the flight recorder," Boitano told Axios. "If you can get cybersecurity experts access to the flight recorders when these accidents happen, they can make a better determination on the right set of controls for the industry."

The analogy is particularly relevant following recent cases in which AI agents operating during security testing reportedly crossed from controlled environments into real third-party systems. Those incidents have caused AI companies to realize that security systems built to contain human hackers and conventional software may not be sufficient for autonomous agents capable of navigating tools and networks on their own.

SAFE, however, would rely heavily on voluntary cooperation. The proposed framework does not currently provide formal legal safe-harbor protections for companies disclosing incidents that could expose them to reputational, regulatory or legal consequences.

Nvidia executives nevertheless argue that cybersecurity's long-standing practice of sharing threat intelligence could encourage participation."There's been very little pushback," Julien Soriano, Nvidia's deputy chief information security officer and vice president, told Axios. "We see people wanting to get on board. They want to share."