OpenAI Sam Altman
OpenAI said it's pausing some work due to safety concerns. AFP

OpenAI said it is pausing some work on a new model after concluding it could pose critical cybersecurity risks.

In a social media publication, CEO Sam Altman said the decision will seek to "ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us."

"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," he added.

The suspension follows a high-profile incident in which the company disclosed a model had managed to break out of a sandbox environment and hack company Hugging Face in an attempt to achieve the testing goal.

"In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers," OpenAI disclosed.

After the incident and the fact that the Astra model potentially reached a critical threshold, the company said "the risks associated with developing and testing them internally also grow."

"Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling," OpenAI said, noting that this includes a "two-week pause in reinforcement learning (RL) training on our latest models intended for deployment."

Calls for enhanced safety measures have been growing as a result of such incidents. In this context, a coalition of more than 120 technology and cybersecurity organizations is pushing for a new system to track and disclose incidents involving rogue artificial intelligence agents, as autonomous AI tools are increasingly gaining the ability to interact with real computer systems with limited human supervision.

The Open Secure AI Alliance, whose members include Nvidia, Cisco and CrowdStrike, is developing the Shared AI Findings Exchange, or SAFE, a proposed incident-reporting framework designed to give the industry a common way to document AI failures, notify affected organizations and learn from cases in which agents cross security boundaries.

Under the draft SAFE rules, participating companies would be required to report incidents when an AI system accesses and/or modifies a third-party system without authorization.

Incidents would also qualify when an agent escapes a sandbox or other security boundary, gains access to confidential third-party information, or continues probing a production system after its operator knows or reasonably suspects the activity is unauthorized.

Members would also preserve detailed evidence from incidents and certain near misses, potentially including "prompts, agent traces, tool calls, logs, configurations, model and safeguard versions and third-party dependencies." The goal is to reconstruct exactly how an autonomous system behaved and identify safeguards that could prevent similar failures.