OpenAI will begin inserting an invisible mark into eligible texts from ChatGPT and Codex in the European Union, turning the AI Act's transparency requirement into one of the largest practical tests ever conducted with text watermarking. The problem appears in the company's own numbers: changing just 10% of the words already reduces the ability to detect the mark from about 92% to 66%. With 25% of the text modified, the rate falls to 17%.
The contrast does not mean that OpenAI is necessarily offering a solution incapable of complying with the legislation. The European rule contains an important caveat: marking systems must be effective, interoperable, robust and reliable to the extent that this is technically feasible, considering the limitations of the content and the state of the art.
The question becomes less whether the watermark can identify all AI-generated text and more what level of reliability will be sufficient to characterize regulatory compliance when the signal itself can disappear after relatively small changes.
Europe requires identification, but not infallible proof
Since August 2, Article 50 of the AI Act requires that providers of systems capable of generating text, image, audio or video adopt mechanisms that make it possible to identify artificial content in machine-readable format.
It is exactly this layer that OpenAI tries to address with textGrain, a technology that statistically modifies the word choices made by the model without inserting hidden characters or visible marks.
The distinction matters because the legislation recognizes the technical limitations of the process. There is no requirement in the legal text that any generated content remain detectable after any possible transformation.
This prevents a simple reading that OpenAI's tests demonstrate a direct incompatibility between the technology and the law. They demonstrate, however, that regulatory compliance and reliable identification of a text's origin are different problems.
A provider can implement a technically adequate mechanism at the time of generation and still be unable to recover the signal after the text circulates, is summarized, translated or rewritten.
Minimal editing already drastically weakens the signal
It is at this point that the numbers released by OpenAI itself become relevant.
In passages of approximately 400 tokens, the company says it detects about 92% of marked texts in a given test configuration. Replacing 10% of the words with synonyms reduces the result to 66%. With 25% of the words replaced, only 17% continue to be detected.

The mechanism explains the fragility. textGrain depends on statistical patterns distributed across the model's linguistic choices. Rewriting the text means replacing precisely part of the elements that carry this pattern.
This creates a particularly important limitation for text because editing is not an exceptional scenario. Human revisions, grammar tools, translation, summarization and rephrasing are part of the normal cycle of content production and distribution.
The difficulty increases in short texts or in content where the model has less freedom of choice. In the tests released by OpenAI, approximately 80% of 200-token texts were identified, versus about 95% among 400-token passages, considering a target false-positive rate of 1%.
Mathematical content also showed lower results than texts with greater linguistic flexibility. The smaller the space of possible choices for the model, the smaller the space available for inserting a robust statistical signal tends to be.
The same watermark does not work equally in all languages
There is also a specifically European difficulty: language.
In tests conducted by OpenAI in the 24 official languages of the European Union, initial detection ranged from 69% in Spanish to 42.2% in Romanian, again considering a target false-positive rate of 1%.
The company says it can strengthen the signal in languages that show lower performance. But the difference between the results reveals another obstacle to a regulatory implementation on a continental scale.
The same mechanism does not necessarily produce the same reliability across the entire European market.
This matters because the regulatory obligation is common, while the statistical properties of languages are not. A technology considered sufficiently effective in English or Spanish may require additional adjustments to achieve similar performance in other languages.
In practice, assessing compliance may depend not only on the existence of a watermark, but also on how it works in different languages, text lengths and content types.
A detected watermark is not equivalent to proving authorship
OpenAI also avoids presenting its detector as a definitive test of artificial content.
The presence of the mark may indicate that an OpenAI system generated or processed part of the text, but it does not determine how much human work was done afterward, who wrote or published the material, who owns the rights to it, or whether the content is true.
The reverse is even more important: not finding the watermark does not allow concluding that a text was written by a person.
The content may have been edited, translated, shortened, rewritten by another model, produced before the implementation of the technology or simply be too short to carry a sufficiently strong statistical signal.

This limitation helps explain why the detector will not be made immediately available to the public. OpenAI intends to restrict initial access to researchers and specialized organizations.
This choice reduces the risk that a probabilistic tool will be turned into an automatic verdict on authorship, especially in environments such as schools, universities, companies or legal proceedings.
Watermarking can produce evidence about provenance. It cannot, on its own, reconstruct the entire creation history of a text.
Provenance tends to work as a signal, not as a verdict
The European design also points in this direction.
In addition to the technical obligation imposed on providers, the AI Act establishes specific rules for those who publish texts generated or manipulated by AI on matters of public interest. This requirement has an exception when there is human review, editorial control and a natural or legal person assumes responsibility for the content.
This creates two distinct layers. One tries to technically preserve a signal about the origin of the material. The other determines who assumes responsibility for its publication.
The watermark does not replace this second layer. OpenAI itself states that invisible marks and provenance systems do not eliminate possible visible disclosure obligations.
The consequence is that text provenance tends to work better as complementary evidence than as an isolated mechanism for deciding whether a given piece of content came from AI or not.
This model already appears in other provenance strategies. For images, for example, companies can combine content credentials, metadata and technical signals because any individual mechanism can disappear during editing, conversion or publication on certain platforms.
In text, the problem is even more difficult. Copying, summarizing, translating and rewriting are not sophisticated attacks against the system. They are common operations.
The regulatory problem becomes defining what is reliable enough
It is precisely there that OpenAI's implementation may gain importance beyond the company's own products.
The AI Act does not require perfect technology. It conditions requirements such as robustness and reliability on what is technically feasible and on the available state of the art.
This protects the regulation from an impossible requirement, but it also opens a practical question: what performance will be considered sufficient for a company to demonstrate that it actually took effective identification measures?
A high rate in intact texts may be sufficient to demonstrate that the system inserts a detectable signal. But it does not necessarily solve the problem that motivated the rule if this signal disappears during common uses of the content.
On the other hand, requiring a watermark to survive any form of transformation would also be technically difficult and could produce other problems, such as false positives or excessive interference in text generation.
The regulatory discussion therefore tends to shift from the existence of the mechanism to its concrete performance.
False-positive rates, resistance to editing, differences between languages, minimum required length and behavior after translation or processing by other models may become more relevant criteria than simply asking whether a given system has a watermark or not.
The decisive test begins after launch
The next relevant indicators will not come only from controlled benchmarks.
The implementation in ChatGPT and Codex will make it possible to observe how textGrain behaves when texts pass through editors, translators, summarization tools, competing models and human review processes.
It will also be important to monitor how European authorities interpret the requirement for robust and reliable solutions to the extent that they are technically feasible. The Commission considers the Code of Practice on transparency a tool to help companies demonstrate compliance, but adherence to the code does not, by itself, represent definitive proof of compliance with the legislation.
Another important signal will be the reaction of competitors. If different providers adopt incompatible identification systems, the interoperability required by the European regime may become an additional problem. If shared standards advance, provenance may evolve from proprietary solutions to a broader identification infrastructure.
OpenAI's experience therefore places a concrete limit on the regulatory discussion. Europe can already require that large models leave technical signs of their participation in the creation of a text. There is still no evidence, however, that these signs can reliably survive the normal cycle of language editing and circulation.
The success of this policy will depend less on turning watermarking into absolute proof of authorship and more on defining when an imperfect signal is still robust enough to be useful, auditable and legally relevant.



