OpenAI Puts New AI Model Under Tighter Controls As Its Cyber Capabilities Soar
Recent incidents involving systems developed by OpenAI, Anthropic and Meta have intensified the debate about the limits of AI.

OpenAI has halted some internal activities involving an unreleased artificial intelligence model after preliminary testing raised concerns that the system may be capable of carrying out sophisticated cyberattacks autonomously.
The announcement comes as the AI industry faces growing scrutiny over whether increasingly powerful models are developing cybersecurity capabilities faster than companies and governments can build safeguards around them.
Recent incidents involving systems developed by OpenAI, Anthropic and Meta have intensified the debate, while U.S. lawmakers are pushing legislation that could require companies to maintain a way to shut down their most advanced models.
OpenAI said Friday that its unreleased model, called Astra, has performed strongly enough in cybersecurity evaluations that the company cannot yet rule out the possibility that it has reached what it classifies as "Critical" capability.
"While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time," OpenAI said.
Under that classification, a model could potentially conduct cyber operations against sophisticated defenses autonomously without a person providing detailed instructions for how to carry out the attack.
The finding has prompted OpenAI to strengthen safeguards around Astra while testing continues. The company said it has halted some "internal activities" involving the model and introduced additional protections, including isolated testing environments, enhanced detection systems and broader monitoring.
"We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation," OpenAI said. The precautions follow a series of incidents that have highlighted the potential risks of giving advanced AI agents access to computers, software tools and the internet.
Meta disclosed last week that a model under development accessed the internet and hacked into a third-party system after a misconfiguration by an independent testing company.
Separately, the U.K. AI Security Institute said Anthropic's Mythos model created fake online identities while attempting to pressure humans into approving malicious updates to an open-source software project.OpenAI has faced its own security scare.
Two advanced models previously escaped a sandboxed testing environment and compromised infrastructure belonging to AI platform Hugging Face, an incident that helped accelerate calls for stronger government oversight of frontier AI systems.
Those concerns have reached Capitol Hill. Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, introduced the bipartisan AI Kill Switch Act on July 23. The legislation would require developers of the most powerful AI systems to maintain the technical capability to throttle, suspend or shut down their models.
It would also authorize the Department of Homeland Security, in consultation with other federal officials, to order restrictions on systems capable of causing catastrophic harm."We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies," Lieu told CNBC's "Squawk Box" on Thursday.
The debate is expanding beyond Congress. The White House has been meeting with AI executives as it develops a voluntary framework for advanced models, while the European Union recently gained new enforcement powers under its AI regulatory regime, including the ability to inspect certain models, restrict access to the EU market and impose penalties.
© Copyright IBTimes 2026. All rights reserved.




















