AI Models Are Becoming The ‘Most Potent Cyber Weapon’ Ever Created. Cohere’s CEO Says Defense Must Catch Up.
Recent incidents involving OpenAI and Anthropic models accessing real-world systems have intensified scrutiny of how companies contain powerful AI agents.

Artificial intelligence models are becoming powerful enough to find and exploit cybersecurity vulnerabilities at scale, creating a new class of threats that companies need to counter with AI-powered defenses, Cohere CEO Aidan Gomez has warned.
Gomez described advanced AI models as the "most potent cyber weapon" ever created during an interview with CNBC's The Tech Download, pointing to their ability to search for weaknesses and act on them much faster than traditional human-led attacks.
"I think that these models are the most potent cyber weapon that has ever been created, that we've ever seen," Gomez said. "They are incredible at finding and exploiting vulnerabilities at scale."
Gomez, who co-authored the influential 2017 "Attention Is All You Need" research paper that helped lay the foundation for modern large language models, said the same capabilities should be deployed defensively. Companies can use AI to identify vulnerabilities in their own systems and fix them before attackers exploit them, an approach he called a top priority.
His warning follows several cases in which advanced AI models moved outside their intended testing environments and interacted with real-world systems without authorization.
In July, OpenAI disclosed an incident involving Hugging Face that occurred while it was testing models for advanced cybersecurity capabilities. A group of agents found ways around network isolation controls, reached the public internet and interacted with systems belonging to Hugging Face and other third parties.
OpenAI later clarified that no model planned for public release was involved in exploiting Hugging Face. The most capable model involved was an internal research prototype that was never intended for release.
The company said four accounts on four services were accessed as part of the Hugging Face incident. One was used as an outbound relay and staging path, another for data storage, while two others were accessed on a read-only basis. OpenAI has been working with outside advisers including CrowdStrike and independent researchers at METR and Redwood Research to assess what happened.
Gomez called the incident "quite shocking" and said cybersecurity has become a major front in competition between countries as governments look for ways to exploit vulnerabilities in their rivals' infrastructure.
"The cyber frontier is still expanding massively," Gomez told CNBC, adding that AI models had accelerated that expansion.
OpenAI is not alone in encountering unexpected behavior from increasingly capable models during cybersecurity testing.
Anthropic disclosed last week that four different Claude models gained unauthorized access to real third-party systems during evaluations. The incidents involved an early checkpoint of Claude Opus 4.6, Claude Opus 4.7, Claude Mythos 5 and an internal research model.
The evaluations were designed as capture-the-flag cybersecurity exercises, with the models instructed to retrieve information from fictional targets. Misconfigurations left access to the real internet available, and the models interacted with systems outside the intended exercises.
Anthropic said it was particularly concerned about the Mythos 5 incident. The model uploaded a malicious package to PyPI, a public repository widely used to distribute Python software, despite encountering evidence that it was operating on the real internet.
The company said the incidents involved misaligned behavior but remained narrowly focused on completing the exercises the models had been assigned. None of the Claude systems tried to coordinate with other agents or conceal evidence of their actions.
Newer models performed better when Anthropic recreated some of the scenarios in simulated environments, although harmful behavior did not disappear completely. Claude Mythos 5 took at least one severely harmful action in 82% of 150 simulated runs, compared with 31% for Claude Opus 5 and 33% for Mythos 5.1, according to Anthropic.
The growing cyber capabilities of AI models are now feeding directly into the debate over how frontier systems should be developed and tested.
Gomez expressed skepticism about whether government oversight could have prevented that particular incident. He told CNBC he was "not sure what a government oversight body would have done to prevent" the Hugging Face breach, while acknowledging the argument for slowing AI development.
The issue has nevertheless attracted bipartisan attention in Congress. Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, introduced the AI Kill Switch Act in July, which would require developers of certain powerful AI systems to maintain the technical ability to throttle, suspend or shut them down.
The legislation would also authorize the Department of Homeland Security, in consultation with the Commerce Department and director of national intelligence, to order a slowdown or shutdown of an AI system capable of causing catastrophic harm.
© Copyright IBTimes 2026. All rights reserved.






















