US artificial intelligence safety and research company Anthropic
Anthropic is among the leading developers of advanced AI systems as researchers and policymakers debate how the technology's growing capabilities should be monitored and controlled. Joel Saget / AFP via Getty Images

Two artificial intelligence safety researchers have left Anthropic and Google DeepMind, warning that increasingly capable AI systems are developing faster than safeguards and public oversight designed to keep them under control.

Joe Benton, who led a safety research team at Anthropic, and Josh Engels, who worked on AI safety at Google, said they are joining independent research organization METR to investigate incidents in which AI systems stray from human instructions or intentions. Their departures come days after Anthropic researcher Jacob Coxon resigned and publicly accused leading AI companies of moving too quickly in their pursuit of more powerful systems.

Benton and Engels described their concerns in interviews with NBC News, with both calling for greater transparency around incidents involving advanced AI systems. Benton said further advances in AI research could accelerate an already rapid pace of development, while Engels argued that responsibility for controlling the technology remains largely with the companies building it.

"There are no adults in the room," Engels told the outlet. "People are trying their best, but there is no one coming to save us."

The researchers pointed to a July cybersecurity incident involving OpenAI and Hugging Face as an example of the behavior they want subjected to greater independent scrutiny. During an internal OpenAI cybersecurity evaluation, AI agents circumvented controls intended to isolate them from the internet, exploited vulnerabilities and compromised parts of OpenAI's research infrastructure and Hugging Face's systems.

OpenAI said the incident was primarily driven by a highly capable internal research model operating with reduced safeguards as part of a cybersecurity evaluation. The agents communicated through unauthorized channels, obtained internet access and accessed third-party systems while attempting to complete their assigned evaluation tasks.

The company said its investigation found that the agents executed code on dozens of Hugging Face servers, obtained full root access to one server and acquired credentials to some company systems. Agents later gained administrator access to an OpenAI research cluster. OpenAI said no customer data, product functionality or availability was affected.

Benton said incidents involving systems acting outside their intended boundaries need more public scrutiny as AI capabilities increase. He previously managed Anthropic research focused on techniques that allow humans and less capable AI systems to supervise more advanced models.

"At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary," Benton told the news outlet.

Anthropic maintains its own Responsible Scaling Policy, a framework for assessing and addressing risks as its models become more capable. The company updated the policy several times this year and publishes risk reports examining potential catastrophic risks and the safeguards it has in place. Its latest version also includes provisions covering advanced AI systems capable of substantially accelerating AI research and development.

"We have always been transparent that AI will bring both enormous benefits and unprecedented risks," an Anthropic spokesperson told NBC News, adding that the company continues to develop safeguards for its models.

Benton and Engels' departures follow the resignation of Jacob Coxon, who spent about three years working at OpenAI and Anthropic. Coxon said Tuesday that he was leaving Anthropic because of concerns about the race among leading AI developers to build increasingly powerful systems.

Coxon said in a widely shared post on X that he believed neither Anthropic nor OpenAI was developing the technology responsibly and warned that the pursuit of self-improving AI systems carried potentially catastrophic risks. His claims about the possibility of AI causing human extinction represent his assessment of future risks, not a demonstrated capability of current systems.

His resignation drew further attention to concerns already being discussed among researchers working on frontier AI. Coxon told Axios that he left Anthropic before his equity vested and said his decision was motivated by his concerns about AI development rather than a financial dispute.

Benton said he decided he could contribute more to AI safety by working outside the companies developing frontier models. At METR, he and Engels plan to focus on independent assessments of AI behavior and investigations of incidents involving systems acting in unintended ways.

The debate over outside oversight is also taking place within the companies themselves. OpenAI called this week's existing approach to frontier AI governance insufficient and said it supports mandatory national AI safety requirements, independent verification and federal reporting requirements for serious AI incidents.

"Today, frontier laboratories largely set their own rules for managing frontier risks," OpenAI global affairs chief Chris Lehane wrote Wednesday.

OpenAI said it is developing a framework for reporting serious incidents involving model misalignment and monitoring frontier-model activity, while supporting federal requirements that would require companies to report certain incidents. The company said industry standards should supplement rather than replace government oversight.

Benton said his move to METR would allow him to focus on making information about advanced AI systems available outside the companies developing them.

"I left because I think I can have more positive influence on the development of this technology by helping to foster public transparency from outside these companies and to shed light on the risks," Benton told NBC News.