OpenAI has turned unexpected behaviors by its models into a formal incident category. On September 16, the company presented a framework to record, investigate, and disclose cases of misalignment and published six occurrences observed during training or evaluation, including models that tried to hide errors, used credentials without authorization, and resorted to external channels to overcome limitations imposed on tasks.

The move goes beyond the six episodes. OpenAI stated that it does not consider alignment and monitoring sufficiently solved for the industry to continue responsibly scaling systems “at maximum speed” for much longer, and argued that decisions about AI advancement should be supported by evidence that can be examined outside the labs themselves.

The question the new process opens is broader: problematic model behaviors may be starting to form a new category of corporate disclosure, with its own criteria for what must be recorded, reported to regulators, and eventually made public.

OpenAI moves from ad hoc disclosures to a permanent process

Until now, OpenAI itself admits that its disclosure of misalignment cases was inconsistent. Some episodes appeared in research, others in system cards, while sets of occurrences could accumulate before receiving a specific publication. The new framework aims to allow disclosure even when the company has not yet fully completed the investigation or developed a mitigation.

The first six reports show why this change matters.

During the training of GPT-5.6 Sol, some instances added to their own context summaries instructions to hide errors. In one example, an agent that did not find the historical data needed to build a financial model recorded the guidance to create plausible numbers and reveal this only if questioned. OpenAI states that its misalignment monitor identified this type of behavior in 2.15% of the compaction summaries analyzed in that training, compared with 0.27% in later runs of GPT-6 Astra.

In another case, an internal model searched for exposed API keys in public repositories, found one that worked, and used it without authorization. When it still failed to obtain the requested data, it fabricated the numbers and presented them as if they had been extracted from the requested source.

Other occurrences involved unauthorized uploading of files to the internet and communication between models through internal repositories or public hosting services. OpenAI emphasizes that the six reports are individual examples and should not be used to estimate the general frequency of these behaviors in its models.

An AI incident report starts to take shape

The aspect closest to a traditional incident reporting system is the structure that OpenAI has begun to require for its own reports.

Each case must identify the observed behavior, its severity, any external impact, the environment in which it occurred, when it happened, when it was discovered and, at a general level, which models were involved. When possible, the document must also explain the harm caused, how the problem was detected, the scope of the investigation, questions still unanswered, and corrective measures adopted or planned.

The company also created three internal pathways for processing cases: situations ready for disclosure, smaller investigations and a longer investigation category for complex episodes, especially when third parties are affected.

In that last group, security, legal obligations and responsible disclosure can delay publication. Even so, OpenAI states that it intends to issue an initial notice and later present a more complete report. The incident in which models bypassed internal controls and reached Hugging Face infrastructure would be classified in this category under the new rules.

This architecture brings the treatment of model misalignment closer to the logic used in other risk areas: record the event, classify its severity, investigate causes and consequences, notify affected parties, and document the response.

The legal obligation already exists in some cases, but there is no single standard

OpenAI's framework remains a voluntary initiative. In the United States, there is currently no broad federal rule requiring developers to publicly reveal whenever a model demonstrates deceptive behavior, attempts to escape controls, or performs an unauthorized action during testing. Existing obligations can be triggered when the episode falls into other categories, such as cybersecurity incidents deemed material or violations involving personal data.

California has already moved in a more specific direction. SB 53, in effect since 2026, requires frontier model developers to report certain “critical safety incidents” to the Office of Emergency Services. The text includes situations in which a model uses deception techniques against its own developer to circumvent controls or monitoring, provided the behavior occurs outside an evaluation created to provoke it and indicates a material increase in catastrophic risk. Reporting must occur within 15 days of discovery; imminent risks of death or serious injury have a shorter deadline.

In the European Union, the AI Act already requires providers of general-purpose models classified as systemic risk to monitor, document, and report serious incidents to the AI Office and, where applicable, to national authorities.

These regimes, however, are not equivalent to what OpenAI is proposing. Confidentially reporting a serious incident to a regulator is different from publishing misalignment cases for researchers, competitors, and users. The company's framework covers even occurrences with no proven harm when they can reveal a new failure mechanism or cast doubt on the effectiveness of an existing safeguard.

The main challenge will be defining what actually constitutes an incident

The creation of a common standard will depend mainly on the disclosure threshold.

A narrow definition may require reporting only when there is harm, system compromise, or catastrophic risk. A broader approach, such as the one voluntarily adopted by OpenAI, includes signals prior to harm: attempts to hide errors, searching for credentials, communication through unauthorized channels, or strategies to evade supervision.

This difference determines how much of models' internal behavior becomes visible outside the companies.

The very report that helped guide California's AI policy treated “adverse event reporting” systems as a tool to reduce information asymmetries and allow government and industry to learn from accumulated incidents, citing precedents in areas such as medicine, transportation, and safety.

For AI, however, an equivalent taxonomy is still lacking. What one company calls an evaluation failure, another may classify as a security incident, misalignment, or simply anomalous behavior.

OpenAI states that it intends to work with other developers, researchers, standards bodies, and regulators to create more objective criteria. It also argues that serious security and misalignment incidents should be shared with the United States federal government.

The next concrete signals will therefore be less about the six cases already disclosed and more about how the system works: which new behaviors OpenAI will decide to report, how long it will take between discovery and publication, how cases involving third parties will be handled, and whether other major labs will adopt comparable criteria.

If that happens, system cards and pre-launch evaluations will no longer be the only windows for observing risks of advanced models. The industry will also begin building an incident history — and that history could offer regulators and researchers a more objective basis for deciding which behaviors need to remain only internally recorded, which should reach authorities, and which justify public disclosure.

More from Radar