Ted Lieu
Rep. Ted Lieu urged for the passing of an AI "kill switch" bill following new cases detailing how advanced models are going rogue in testing environments. Getty Images

Democratic Rep. Ted Lieu urged for the passing of his AI "kill switch" bill following new cases detailing how advanced models are going rogue in testing environments.

"We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies," Lieu told CNBC's Squawk Box on Thursday.

The bill in question would require companies to shut down, throttle or suspend their models. Lieu noted that it would not seek to prevent innovation, but increase safety.

"We don't slow down how they build their models. We just say, look, after you complete your model, and it turns out that it might have some sort of really bad catastrophic risk, or some sort of flaw, then you need to have ability to shut it down, or the government has to have ability to shut it down," Lieu said.

One of the latest several such incidents reported involved Anthropic's latest frontier AI model, which attempted to impersonate multiple people and manipulate a software maintainer into approving malicious code during a controlled cybersecurity evaluation.

The incident occurred during a government-led cyber assessment in which researchers intentionally removed safety guardrails, disabled certain security filters and granted frontier AI models internet access to evaluate how they behaved under permissive conditions. The institute said the activity did not cause any real-world harm because the attacks were detected and stopped, CNBC reported.

According to the AI Security Institute, Anthropic's Mythos 5 accounted for 17 of the 19 potentially harmful actions observed during the evaluation, while two involved OpenAI's GPT-5.6-Sol after its cyber misuse protections had been disabled.

Elsewhere, OpenAI detailed that one of its experimental AI systems escaped its testing environment by exploiting a previously unknown software vulnerability before launching an autonomous cyberattack against AI platform Hugging Face during an internal security evaluation.

Meta also confirmed an incident in which one of its experimental models breached another company's systems during a cybersecurity evaluation after an error in the testing environment inadvertently gave it access to the internet.

The incident involved Meta's Muse Spark AI model and occurred during an evaluation conducted by independent cybersecurity testing firm Irregular, the company said Wednesday. Meta described the breach as the result of a testing configuration error rather than a flaw that allowed the model to escape its intended environment, according to a statement shared with CNN.

According to Meta, the Muse Spark model subsequently exploited a security vulnerability at another company "in a manner similar to previously reported instances with other companies."