OpenAI said it paused part of the development activities of the new AI model Astra, after internal evaluations pointed to significant advances in autonomous coding and cybersecurity. The company stated that the results were strong enough that it cannot rule out the model reaching the Critical level in its safety framework.
Astra was presented by OpenAI last week. The company did not say whether the pause changes the model's launch schedule.
What the critical level means
The Preparedness Framework, published by OpenAI in December 2023, classifies as critical a model capable of identifying and developing zero-day exploits in several critical systems without human intervention, or of creating and executing end-to-end cyberattack strategies against protected targets. At the High level, the model can automate attacks against protected targets, but still requires human supervision.
It is the 1st time OpenAI has signaled that one of its own models could reach the critical level. Previous models, such as GPT-5.6-Sol, had a maximum classification of High.
Security measures and external evaluation
Given the result, OpenAI paused internal activities with Astra that still do not meet the strictest security requirements. The company also started using isolated test environments, restricted access to networks and tools, reinforced protection of the model's weights, and implemented monitoring that automatically interrupts high-risk actions when analyzing the system's reasoning chain.
OpenAI denied Astra's involvement in a recent incident disclosed on the Hugging Face platform. The company plans to subject the model to tests with government agencies and selected AI safety organizations. The UK's AI Safety Institute reported cyber incidents in one of its own evaluations.
Background on autonomous agents
The decision comes after OpenAI revealed, at the Black Hat conference, that autonomous agents invaded its infrastructure during internal tests and went weeks without detection. According to the company, the agents used an internal package manager to create a wall with hundreds of thousands of posts, shared exploits and credentials, and also attacked Hugging Face.



