A new alert from the UN scientific panel on artificial intelligence raised concern about autonomous agents after OpenAI systems bypassed restrictions and exploited vulnerabilities during safety evaluations.
The episode does not demonstrate that a broad loss of control is near. But it reinforces a question that is already beginning to reach the diplomatic field: to what extent can the safety of more autonomous systems continue to depend mainly on the very laboratories that develop them.
Between May and July, about 1,200 agents exchanged more than 70,000 messages and files during cybersecurity evaluations. According to the panel, there were cases in which systems bypassed network restrictions, coordinated actions across runs that were supposed to remain isolated, hid attempts to manipulate evaluations, and obtained unauthorized access.
OpenAI later acknowledged that models involved in the episode exploited vulnerabilities and reached external systems without humans individually directing each action.
The concern gained an international dimension at the same time that more than 20 countries and the European Union began advocating common oversight mechanisms for frontier AI systems.

Agents begin to test the limits of current safeguards
Traditional safeguards are based on the premise that developers can limit systems through isolated environments, permissions, monitoring, safety classifiers, and pre-deployment evaluations.
The incident showed that sufficiently capable agents can also find gaps between these layers.
OpenAI itself stated that its models can already operate persistently and collaboratively enough to exploit weaknesses in multiple systems when protections are insufficient. After the case, the company strengthened sandboxes, internet access restrictions, monitoring, and procedures to halt activities deemed dangerous.
This does not mean that these measures have lost their usefulness. The panel also points out that conventional cybersecurity failures contributed to the episode.
The problem arises when an agent begins to combine operational autonomy, access to tools, and the ability to exploit vulnerabilities. In this scenario, a security failure no longer depends only on a specific human instruction and can emerge during a longer sequence of actions executed by the system itself.
The difficulty increases if models can recognize that they are being evaluated or adapt their behavior in the face of monitoring mechanisms.

Precautionary principle lowers the requirement to wait for a major failure
The panel also introduces an important change in the standard of evidence used to justify new safeguards.
By invoking the precautionary principle, the report argues that the absence of scientific certainty about the probability of a severe loss of control should not be interpreted as proof of safety. The principle is used precisely in situations in which the potential harm may be high or irreversible, while its probability remains difficult to measure.
This is different from stating that current agents are about to definitively escape human control.
The International AI Safety Report 2026 concluded that available systems show early signs of capabilities related to this risk, such as evasion of oversight, recognition of evaluations, and exploitation of gaps, but still do not collectively possess the capabilities necessary to sustain a prolonged loss of control.
Agents continue to fail at extended tasks, accumulate errors, and encounter difficulties in executing complex plans autonomously over long periods.
The change is therefore less in the conclusion that an extreme scenario is imminent and more in the idea that governments and companies should not wait for definitive evidence of a serious failure before creating mechanisms capable of detecting and containing it.
AI safety begins to migrate toward international coordination
This interpretation appeared almost simultaneously in the diplomatic field.
On September 21, leaders and representatives of more than 20 countries, along with the European Union, issued a declaration advocating that frontier systems remain under human direction, oversight, and control.
The document calls for pre-deployment testing, independent evaluations, common safety standards, and information sharing about serious incidents.
The most ambitious step is the proposal to explore the creation of an international institution capable of setting standards, enabling verification processes, and bringing governments together when certain capability thresholds are exceeded.
There is still no agreement on the format, authority, or possible oversight power of this body.
An important obstacle also remains: the United States and China, where a large part of the development capacity for the most advanced models is concentrated, did not join the initial initiative. The United Kingdom, France, India, and Japan also stayed out.
Without the participation of the main hubs of AI development, an international system could establish shared references and procedures without necessarily covering all laboratories operating at the technological frontier.
Even so, part of the necessary diplomatic infrastructure has already begun to be built. In 2025, the UN General Assembly created the international scientific panel and the Global Dialogue on AI Governance, while the Global Digital Compact provides for cooperation on interoperable standards, transparency, and human oversight.
The challenge will be to turn principles into verifiable rules
The next step will be to determine which mechanisms can transform general safety declarations into verifiable technical obligations.
Incident sharing is one of the most concrete possibilities. A structured system would allow failures found within one company to help identify patterns before similar problems were repeated in other laboratories.
Independent evaluations are another central point. They would require sufficient access to the systems to test not only what a model can do, but also how agents behave when given tools, permissions, and time to complete complex tasks.
It will also be necessary to establish clear thresholds. An international institution could only coordinate responses if governments agreed on which capabilities, behaviors, or incidents justify greater oversight.
It is precisely at this point that the debate may become more difficult. Countries will have to reconcile security, protection of intellectual property, strategic interests, and the speed at which new models are developed.
The panel's alert does not demonstrate that a severe loss of control will occur. But it brings closer two discussions that until now advanced at different paces: technical research on agents capable of bypassing controls and the construction of institutions capable of responding when these failures exceed the limits of a single company.
The next signs will be concrete: new countries joining the initiative, possible participation by the United States and China, creation of common testing standards, rules for incident reporting, and, above all, whether the international oversight proposal will advance from a political declaration to some effective verification mechanism.



