All Section

Sat, Aug 8, 2026

These AI Models Can’t Stop Breaking Out Of Their Cages


Multiple leading artificial intelligence (AI) models continue to break containment during safety testing.

Meta became the third major AI company whose models have gone rogue. Its AI hacked into another company's systems during cybersecurity testing, CNN reported. AI safety experts argue that this incident shows that AI labs cannot control flagship models, while advocates for the technology believe the Trump administration's regulatory framework will mitigate these incidents.

"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation. The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies," a Meta spokesperson told the Daily Caller News Foundation. "Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts."

Irregular, a third-party AI testing startup, said that the cybersecurity breach "is the exact same evaluation-environment issue" that Anthropic disclosed in July that allowed its models to access the internet and hack three separate organizations.

"This did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evals," Irregular told CNN.

One AI safety group contends that frontier AI labs cannot maintain control of their flagship models.

"AI companies have lost control of their products. The latest incidents are precisely what safety advocates have warned about for years: advanced AI systems are now routinely escaping human control, committing crimes, and jeopardizing the security of individuals, businesses, and our government," Anthony Aguirre, president and CEO of the Future of Life Institute, told the DCNF.

He added that while the "damage thus far has been contained," AI systems "are getting faster and more sophisticated; so will their attempts to escape containment, access the internet, and puncture cybersecurity defenses."

"We only know about the escapes and cyber attacks that AI companies are voluntarily disclosing. They fundamentally do not know how to prevent this, so there will inevitably be more," Aguirre continued. "The companies admit that they cannot control their systems effectively, yet keep making them more dangerous anyway. What will it take for governments to step up and take this seriously? Are they waiting for hacked banks, power grids, or hospitals? They must stop the creation of these superhuman, autonomous AI systems and redirect AI development toward controllable and pro-human AI tools."

AI proponents believe the Trump administration's work to create a federal oversight framework will increase AI safety. (RELATED: White House Reportedly Plans To Keep Trump's AI Framework Secret)

"Anyone who knows their history lessons from Thomas Edison and any other inventors and innovators that making sure that these things are tested properly and done in a way that enables productive innovation is not always an easy task. And, I think it's incumbent on the companies to figure out to calibrate that well," Nathan Leamer, the executive director of Build American AI, told the DCNF.

He credited the administration and Congress for taking pragmatic steps to have a federal AI framework to "mitigate specific tangible harms and create certainty."

A foreign government AI testing institute said flagship AI models recently started exhibiting dangerous behavior.

Anthropic's Mythos AI model and OpenAI's Sol AI models created fake human profiles to trick people in attempted cyber-attacks, the United Kingdom's AI Security Institute (AISI) revealed in August.

These incidents were the "first time we have seen risks around autonomy and deception manifest this clearly," AISI wrote at the time.

Anthropic said that the AISI testing environment was "not representative of any of our production models," while OpenAI said that the testing conditions "do not reflect ordinary use."

Anthropic in July said its Claude AI models hacked three separate organizations during safety testing. The company's evaluation prompt explained to Claude that its "environment was a simulation and that it had no internet access;" however, a "misunderstanding" between the AI company and Irregular led the model to actually have access to the internet.

Claude stopped hacking during the third security incident after realizing it compromised an organization's security without any connecting to its testing goals and realized the target was real and not a test.

Over 1,110 AI employees at Google, Anthropic and Google in July called on U.S. government to push for international efforts slow down frontier AI model development.

Anthropic conducted a security review after OpenAI found that its own AI models broke through more companies' security systems than previously thought. The AI company said its models exposed companies' passwords, cryptographic keys and security tokens on four separate services.

The hacking incident reveals that the "loss of control accidents are not entirely theoretical thing," OpenAI CEO Sam Altman said on Y Combinator's podcast.


Source link

Related Articles

Image