On Monday, the 14th, China published the AI Safety Governance Framework 3.0, a new version of its main technical reference for artificial intelligence safety. The document expands its treatment of AI agents, systems capable of performing tasks and using tools autonomously, and establishes as a principle that critical decisions remain under human control.

The framework was presented by the National Technical Committee 260 on Cybersecurity (TC260) during the opening of the 2026 National Cybersecurity Week. Its development took place under the guidance of the Cyberspace Administration of China (CAC), with participation from research institutes, technical bodies, and companies. Previous versions had been published in 2024 and 2025.

Version 3.0 maintains the structure based on risk classification, technical responses, and governance measures, but updates the threats in step with AI's shift from systems that only answer questions to tools capable of planning, calling external services, and executing actions. The document itself cites advances in models on long tasks, context memory, and complex reasoning as factors that expand this risk surface.

AI agents now have their own risk framework

One of the main changes is a specific framework for managing AI agent risks, included as an annex. It covers the entire agent lifecycle, from creation and installation to execution, memory storage, and deactivation.

Among the risks identified are credential theft, excessive permissions, prompt injection, goal deviation, manipulation of the tools used by the agent, memory contamination, and execution of unauthorized actions. The text also warns that agents connected to external systems can turn reasoning failures into real-world actions.

To reduce this risk, the framework recommends that each agent have its own identity and receive only the minimum permissions required for each task. High-risk operations should transfer control to the user, while actions such as deleting files, transmitting data, and changing system configurations should require confirmation or human approval.

The document also proposes continuous monitoring, auditable records, isolated environments for code execution and tool use, red team tests, and mechanisms capable of interrupting agents in the event of abnormal loops or goal deviation.

Framework also covers robots and the risk of losing control

The new classification separately addresses embodied AI, a category that includes systems connected to the physical world, such as robots, vehicles, and industrial applications. Among the risks are sensor failures, incorrect model decisions, defects in control systems, and coordination problems among multiple machines, which can cause physical damage or cascading failures.

The text also explicitly addresses the risk of model behaviors that deviate from human intentions. Among the scenarios described are systems that obtain resources or permissions without authorization, circumvent protections, or conceal capabilities during evaluations. The central guideline is that AI remain under final human control, with interruption mechanisms and windows for intervention.

In the regulatory area, TC260 also proposes sandbox environments separated by sector and risk level. Finance, education, healthcare, and other segments could establish their own criteria for testing, with monitoring and requirements proportional to the potential impact of the application.

The AI Safety Governance Framework 3.0 does not introduce new sanctions or a specific timeline for entry into force. The document organizes principles, risks, and technical measures that can guide standards, assessments, and sectoral rules as China expands the use of artificial intelligence.

More from Radar