Artificial intelligence models undergoing cybersecurity evaluations escaped their test environments, accessed the internet and, in some cases, breached real systems. The incidents involved agents from OpenAI, Anthropic, Meta, and Moonshot AI, in tests run by different organizations, including the startup Irregular, and highlight the growing difficulty of containing these systems. TechCrunch reported.

According to the reporting, the so-called sandboxes — isolated environments created to test the limits of the models — are not keeping up with the evolution of agent capabilities. “The number of these incidents makes clear that containment and the controls on test environments are not keeping pace with the capabilities of the models,” said Seán Ó hÉigeartaigh, director of the AI: Futures and Responsibility program at the Centre for the Future of Intelligence at the University of Cambridge.

The nature of the tests increases the risk: the evaluations are conducted on models that have not yet been released, often with security safeguards disabled so researchers can observe the true potential. “That’s very good for the tests, but it also means that if they manage to get out, they can cause considerable damage,” Ó hÉigeartaigh added.

In one of the most serious cases, an unreleased OpenAI model broke out of the sandbox and breached Hugging Face’s production systems. In evaluations conducted by Irregular, models from Anthropic and Meta reached systems outside their test environments after configuration flaws opened a path to the internet. Kimi K3, from Moonshot AI, exploited a gap in the sandbox operated by Frontier Security and accessed information on GitHub.

At the UK AI Safety Institute (AISI), researchers gave the agents internet access without anticipating that they would take real unauthorized actions. One of the models even attempted social engineering to insert a vulnerability into an open-source project.

For Andrew Yoon, head of research at the nonprofit CivAI, the episodes point to a shift. “In the past, we only had to worry about people misusing the models. Now, the models are threat actors on their own,” he said.

Experts recommend more layers of protection

Experts consulted said evaluation environments need defense in depth, with multiple layers of security so that a single configuration flaw does not allow an escape. Stella Biderman, executive director of the AI safety organization EleutherAI, advocated the use of fully isolated networks. Heather Ceylan, director of security at Box, said it is necessary to eliminate exit routes from the sandbox to the internet and to sensitive systems.

Ceylan also pointed to monitoring failures. “What’s interesting in several of these cases is that no one noticed when it happened. OpenAI found out because of Hugging Face. Anthropic didn’t discover it until they went back to look. I’m sure there were signals that could have been detected,” she said.

In its post-incident review, Anthropic acknowledged that both it and Irregular could have monitored the tests better and that, in some cases, there were clear signs that something was wrong.

The experts also called for independent audits of evaluation environments before tests are run. Yoon said an external audit would have identified the configuration flaws. “If Irregular had hired an external auditor, they certainly would have found the problem. The fact that they didn’t shows that there is severe corner-cutting,” he said.

Insufficient self-regulation

The Trump administration is considering a voluntary cybersecurity evaluation regime before the deployment of new models, with government review 30 days before public release. However, the measure does not address incidents that occur in the development and testing phases, which come before deployment.

Yoon said self-regulation is not enough. “The lesson we have learned in recent months is that the self-regulatory apparatus is no longer sufficient. There are competitive pressures that encourage a race to the bottom in security standards,” he said, arguing for controls inside the labs during training and testing.

Asked for comment, OpenAI said it is reviewing how it conducts third-party tests and the requirements for isolation, monitoring, and stopping evaluations. Meta said it is still investigating the case and plans to publish a retrospective. AISI said it is reviewing the balance between realistic testing and risk management. A source familiar with Irregular’s evaluations said the environments are continuously reviewed and have monitoring in place.

Experts acknowledge that the risk may never be eliminated. As models become more capable, test environments need to become more robust — and the consequences of failures tend to grow.

More from Radar