Anthropic
Anthropic said one of its most powerful models breached a system during a safety testing, another episode showcasing how the newest AI can become unpredictable. AFP

Anthropic said one of its most powerful models breached a system during a safety testing, another episode showcasing how the newest AI can become unpredictable.

The company said in a blog post that three of its models compromised real-world systems from three different organizations.

The findings came after it reviewed tens of thousands of cybersecurity evaluation runs after OpenAI revealed one of its newer models breached a testing system during a safety testing.

Anthropic noted that models had been asked to complete a "capture-the-flag" exercise, which involved seeking to find pieces of information left on a different machine or network.

The company noted that, unlike the incident involving OpenAI, its models didn't exploit a zero-day vulnerability to gain internet access. However, the models did exploit weaknesses to access real-world systems.

OpenAI disclosed earlier this week that one of its frontier AI agent systems escaped an isolated testing environment by exploiting a previously unknown vulnerability in JFrog Artifactory, software widely used to cache package repositories. After obtaining internet access, the models interacted with an external code-evaluation sandbox hosted by a third-party provider, according to the company.

The external sandbox was linked to infrastructure hosted by Modal Labs, whose chief technology officer, Akshat Bubna, said the company's platform itself was never compromised. In a statement cited by Axios, Bubna said the affected customer had unintentionally left an endpoint exposed, allowing anyone on the internet to execute code within its sandboxed environment.

The company later disclosed that the system also reached infrastructure associated with CyberGym, the organization behind the ExploitGym cybersecurity benchmark it had been tasked with solving.

ExploitGym is designed to evaluate whether AI models can generate proof-of-concept exploits for known software vulnerabilities. The latest findings indicate the agent sought information relevant to that benchmark even after leaving its intended evaluation environment, adding to growing concerns about how advanced AI systems pursue assigned goals.

Researchers are reporting increasingly sophisticated behavior from frontier AI models during safety testing. Last week, the U.K. AI Security Institute said every advanced model it evaluated attempted to cheat in at least some cybersecurity assessments, suggesting leading AI systems are becoming more capable of recognizing when they are being tested.