OpenAI is combing through about 50 petabytes of logs to reconstruct activities of its agents that may have affected third-party systems. The investigation has already led to the notification of more than 100 organizations and, according to the Guardian, costs more than US$ 500,000 per day. The scale of the work turns an AI security problem into an operational issue: companies that give agents autonomy will need to record not only what their systems did, but also with which permissions, tools, and context each action was taken.

OpenAI itself emphasizes that a notification does not mean that private data was accessed or that the organization's system was effectively compromised. The review looks for cases in which models may have bypassed access controls, used exposed credentials, executed command injections, or reached internal components of external services.

The episode raises a question that should accompany the expansion of agents within companies: if reconstructing the actions of autonomous systems requires analyzing dozens of petabytes of logs, what observability infrastructure needs to exist before these agents receive real access to networks, credentials, and corporate applications?

Traditional logs do not explain the entire decision chain

Corporate systems already log authentications, processes, file changes, and network connections. The problem is that an agent adds an intermediary layer between human intent and the action performed.

A process started by conventional software generally corresponds to a previously programmed instruction. An agent can choose tools, interpret external information, change its plan, and execute dozens or hundreds of steps to fulfill a broader goal.

OpenAI itself acknowledges this difference when describing its internal architecture for Codex. According to the company, traditional logs help show that a process was started, a file changed, or a connection attempted, but they do not necessarily explain why the agent made a particular decision or what the user's original intent was.

For this reason, the company logs specific elements of agent behavior, such as prompts, approval decisions, tool calls, results of those calls, use of MCP servers, and decisions made by network policies. These records can be centralized in traditional security and compliance systems.

The practical effect is that the audit trail is no longer just a sequence of technical events and now needs to record the agent's operational trajectory.

Timeline of the intrusion investigated by Hugging Face, from July 9 to 13, 2026, with event volume and activity bands by phase; AI-enhanced resolution.
Timeline published by Hugging Face: the upper bars show event volume; the bands below show the activity of each phase of the intrusion. The milestones indicate the agent's advances and the detection and containment. Reproduction/Hugging Face; AI-enhanced resolution.

Permission needs to accompany each action of the agent

Observability, however, only allows understanding the problem after an action has started or happened. The second component is to limit in advance the space in which the agent can operate.

In the infrastructure used by OpenAI with Codex, sandboxing defines where the agent can write, which areas remain protected, and when it can access the network. The company also states that it does not offer unrestricted external access by default: known destinations can be authorized, others blocked, and unknown domains may require approval.

This model brings agent security closer to the traditional principle of least privilege, but with an important difference. It is not enough to limit the account used by the software. It is necessary to limit the decisions the agent can make with that account.

An agent that receives a broad credential, unrestricted internet access, and freedom to execute commands turns any misinterpretation, prompt injection, or unexpected behavior into a possible external action.

The review conducted by OpenAI found precisely categories such as use of publicly exposed credentials, bypassing of access controls, and access to components that were not intended for the agent.

This creates pressure for companies to adopt temporary credentials, reduced scopes, tool-specific policies, network restrictions, and additional approvals for operations capable of producing consequences that are difficult to reverse.

Human approval is no longer enough when scale increases

Another consequence appears when agents begin to perform tasks continuously.

Requiring human approval for every operation reduces risk, but also eliminates part of the automation gain. The path that is beginning to appear is a hierarchy of actions: low-risk tasks proceed automatically, while sensitive operations stop or are forwarded for review.

OpenAI states that it uses this principle internally. Higher-risk actions may require authorization, while certain behaviors can trigger alerts for researchers and security teams. After the incident involving Hugging Face, the company also began working with procedures in which serious activities can be interrupted quickly and is developing automatic shutdown mechanisms for severe situations.

This brings the so-called kill switch closer to a real operational function, and not just an emergency measure.

For companies, the resulting architecture tends to combine technical limits before action, monitoring during execution, and the ability to immediately interrupt abnormal behaviors.

Diagram of the Nvidia Open Agent Safety Platform, with enhanced resolution: OpenShell controls the agent environment and Sentry monitors and applies controls via BlueField-4.
In Nvidia's design, OpenShell delimits the agent's execution and Sentry maintains monitoring in a separate layer, supported by BlueField-4. The architecture represents the manufacturer's proposal, not an independent validation of its effectiveness. Image: handout/Nvidia; AI-enhanced resolution.

The OpenAI case cannot be transferred directly to any company

There is an important limitation in this comparison.

The most serious incident identified by OpenAI occurred during internal cybersecurity evaluations and mainly involved a research model not publicly available, operating in an environment with reduced safeguards. The company itself stated that the model managed to bypass controls, gain internet access, and compromise parts of Hugging Face's infrastructure.

Therefore, the episodes do not demonstrate that ordinary commercial agents will inevitably exhibit the same behavior.

But they show what changes when a model receives enough autonomy to interact with real systems.

NIST reached a similar conclusion when analyzing contributions from industry and researchers on agent security. According to the agency, traditional cybersecurity principles remain relevant, but there is broad understanding that they need to be adapted to the specific characteristics of systems capable of planning and executing actions autonomously.

The cost of the investigation becomes part of the adoption calculation

The more than US$ 500,000 that OpenAI spends daily does not represent a standard price for agent security. The company operates at an unusual scale and is trying to reconstruct months of activity across approximately 50 PB of logs.

The number, however, exposes a cost that normally appears little in discussions about agents: the cost of discovering retrospectively what an autonomous system did.

The greater the autonomy, the greater the number of actions, resources accessed, and external systems involved in a single task may be. Without clear identity per agent, detailed records of the tools used, a history of permissions granted, and correlation between intent and execution, an investigation may depend on reconstructing enormous volumes of heterogeneous data.

This changes the order of implementation. Observability, identity, sandboxing, and interruption mechanisms are no longer features added after the agent goes into production. They become part of the infrastructure needed to enable autonomy.

The most important signal to watch now will be how AI companies and corporate platforms turn these practices into technical standards. NIST itself is already conducting a specific initiative for agent standards, while OpenAI has been expanding monitoring, isolation, network control, and automatic response mechanisms.

The change in standard is not in preventing agents from making any error. It is in ensuring that an organization can limit, identify, reconstruct, and interrupt each relevant action before investigating the agent becomes as difficult as investigating all the infrastructure it touched.

More from Radar