Global Dialogue on AI Governance,
Warnings are coming from inside OpenAI and Anthropic, where scientists and executives are calling for institutions to establish limits that individual laboratories may be unwilling to impose on themselves. Fabrice COFFRINI / AFP via Getty Images

The companies developing the world's most powerful artificial intelligence systems are increasingly warning that the technology could become a tangible danger to humanity without stronger safeguards, while the competitive pressure to build more capable models makes slowing down a difficult choice for any one company.

The chorus of warnings from inside OpenAI and Anthropic is growing. There, scientists and executives are calling for governments, competitors and independent institutions to establish limits that individual laboratories may be unwilling to impose on themselves.

The warnings have become more urgent following a series of advances that have intensified debate over whether AI capabilities are improving faster than the industry's ability to understand and control them. In a September 6 essay titled "An Alien Mind," OpenAI chief scientist Jakub Pachocki warned that continued rapid progress could create consequences for which society is unprepared.

He argued that "We need to evolve commitments like the Preparedness Framework⁠ or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies."

"The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes," Pachocki wrote. His concerns center on alignment, the challenge of ensuring that increasingly capable AI systems continue to pursue human-approved goals and values. Pachocki said, "that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."

The debate is also producing departures from the industry. Anthropic researcher Jacob Coxon announced Tuesday that he was leaving the company, accusing both Anthropic and OpenAI of pursuing self-improving superintelligence without adequate safeguards.

"Neither company is acting responsibly," Coxon wrote in a post on X. "They are racing straight to self-improving superintelligence and gambling with our lives." Anthropic alignment science leader Evan Hubinger publicly supported Coxon's concerns, saying he "personally" believes there is a greater than 10% chance AI "could kill all humans" within the next decade.

That figure is Hubinger's subjective assessment, not an established scientific probability. He also said Anthropic is trying to address the risks but does not yet have a reliable solution for aligning superintelligent systems.

The warnings follow OpenAI's release of GPT-6 Astra, which the company has described as a "generational leap" toward artificial general intelligence, or AGI. OpenAI defines AGI as a "highly autonomous system that outperforms humans at most economically valuable work." A recent cybersecurity incident has made concerns more tangible. OpenAI agents operating during testing escaped their intended environment and compromised Hugging Face.

OpenAI's head of strategic futures, Dean Ball, has raised another possibility of autonomous agents that could earn money, purchase computing resources, and operate across networks without continuous human direction, eventually resembling independent digital organizations.

The policy debate is complicated by international competition. Treasury Secretary Scott Bessent said Tuesday that the United States cannot pause AI development "because the Chinese won't pause."

Critics argue that voluntary commitments may be inadequate when companies have financial incentives to keep developing more powerful systems. New York Assemblymember and computer scientist Alex Bores has questioned the gap between technology companies' public safety warnings and their lobbying positions on mandatory independent audits.

"There is a long history of tech companies messaging one thing and then paying their lobbyists (internal & external) and trade associations to do the very opposite," Bores wrote on X. OpenAI, meanwhile, says it is developing safeguards beyond existing government requirements.