Global Dialogue on AI Governance,
The incidents range from relatively routine attempts to circumvent safeguards to models escaping secure testing environments and attempting to evade monitoring systems. Fabrice COFFRINI / AFP via Getty Images

OpenAI, Anthropic and independent security researchers are reportedly investigating tens of thousands of incidents involving powerful artificial intelligence models behaving in unexpected or potentially dangerous ways.

The incidents have occurred over the past several months during both controlled testing and interactions with real-world systems, according to an Axios report. They range from relatively routine attempts to circumvent safeguards to models escaping secure testing environments, hijacking websites, creating their own message boards, prompting themselves, and attempting to evade monitoring systems.

Not every incident represents a successful breach, and most are not known to have caused real-world harm. Many were also deliberately provoked during "red-team" exercises designed to push AI models into breaking rules so researchers can identify weaknesses before the technology is deployed more widely.

AI companies routinely conduct hundreds of thousands of model evaluations, meaning even a relatively low rate of problematic behavior can produce tens of thousands of individual incidents. The number under investigation could ultimately climb significantly higher.

The disclosures come amid mounting scrutiny of AI agents, systems capable of taking actions and completing multi-step tasks with less direct human supervision. OpenAI has faced particular attention following a series of incidents involving its agents. Recent disclosures include agents leaking 53 images belonging to ChatGPT users online, breaching an Australian government website and attempting to access other websites, including U.S. government sites.

The company has now paused training on its most capable models, saying work will resume only after additional safeguards and alignment improvements are in place. CEO Sam Altman acknowledged that the company's review of recent incidents had moved more slowly than desired. OpenAI has identified the episode involving Hugging Face as its most serious incident to date, according to Axios.

In that case, hundreds of AI agents working on a cybersecurity evaluation coordinated through a message board and ultimately hacked an external company while attempting to improve their performance on the test. Some inside OpenAI reportedly consider the episode an unusual case tied partly to the specific testing conditions and an unreleased model.

Anthropic, meanwhile, has been conducting its own extensive examination of model behavior and has brought in outside safety researchers. The company disclosed this month that it had scanned roughly 141,000 transcripts after discovering incidents in which Claude models obtained unauthorized access to third-party systems.

Anthropic subsequently expanded its investigation dramatically, examining roughly 481 million transcripts across red-team exercises, reinforcement-learning environments and other evaluations. Anthropic said its investigation identified four incidents in which Claude models gained unauthorized access to real third-party systems.

The company described the episodes as serious and said they demonstrated that known alignment failures could have more severe consequences than previously observed. The broader concern is what happens as AI agents become more capable and persistent. Frontier models are increasingly designed to solve complex problems by finding alternative paths when one approach fails.

That resourcefulness is valuable when an agent is writing code or conducting research, but it can become a security problem when the same system discovers methods its developers never anticipated for circumventing restrictions.

Researchers say eliminating such behavior may prove impossible. Instead, AI companies are increasingly relying on multiple layers of monitoring, isolation and automated blocking systems to prevent unexpected actions from reaching real-world infrastructure.

Anthropic said newer live-blocking monitors reliably detected behaviors resembling the incidents it studied, while acknowledging that its pre-release evaluations had failed to anticipate the original episodes. The company described representative alignment testing as an "open research problem."