EU AI Act official banner

OpenAI is putting an invisible fingerprint on its text output. On October 5, 2026 (UTC), the company announced textGrain, a statistical watermarking technology that embeds an undetectable signal into the words ChatGPT and Codex choose as they write. The rollout is tied directly to the European Union's AI Act, which requires generative AI providers to make machine-detectable content provenance available — and it signals that the era of unmarked AI text is ending.

How textGrain works

The technology is clever in its simplicity. Instead of appending hidden characters, invisible spaces, or unusual punctuation — approaches that are easy to strip out — textGrain modifies the model's own word selection process. As the model generates text, it uses a cryptographic key and the surrounding context to divide its vocabulary into blocks. An optimal-transport algorithm with KL regularization then ties which block gets selected to a key-generated random number, while preserving the original relative probabilities of words within each block.

The result is a statistical pattern woven into the wording itself. A detector with the key can score a passage and determine whether it carries the OpenAI watermark, even after some editing. Readers see nothing different — the text reads naturally, and benchmark performance barely moves. In OpenAI's testing on its Astra model, GPQA Diamond scores went from 94.44% without watermarking to 93.94% with it — a difference within noise.

The rollout plan

OpenAI is taking a phased approach, and the API gets the first look:

OpenAI also announced its support for the European Commission's Code of Practice on Transparency of AI-Generated Content, joining Anthropic, Google, Meta, Microsoft, and Mistral. xAI notably did not sign.

What the watermark can and can't do

OpenAI is being unusually transparent about textGrain's limitations, and that honesty matters. Detection accuracy scales with text length: at a 1% target false-positive rate, a 200-word paragraph is detected about 80% of the time, while a 400-word paragraph jumps to about 95%. Mathematical content performs worse because the constrained vocabulary leaves less room for the statistical signal to hide.

Editing erodes the watermark. Replacing 10% of words with synonyms drops detection from roughly 92% to 66%. Replace 25%, and detection falls to 17%. That means a determined user can strip the signal with a paraphrasing pass — a reality OpenAI acknowledges rather than paper over.

The detector itself is not publicly available. Access is limited to approved researchers and professional institutions, a decision that balances utility against the risk of adversaries reverse-engineering evasion techniques. OpenAI says it plans to open-source the technology in the coming weeks, which will let the broader research community test, stress, and potentially improve it.

Why this matters

The EU AI Act's transparency requirements are the first binding legal mandate for AI content provenance at this scale, and OpenAI's implementation is the template every other provider will be measured against. If textGrain works well enough to be useful but weak enough to be trivially defeated, regulators may push for stronger mandates — and the industry may have to adopt more invasive watermarking methods that do affect output quality.

The opt-in API default is also telling. OpenAI is not forcing watermarking on developers outside the EU, which suggests the company sees a tradeoff between compliance and user experience. Enterprise customers who need provenance for regulatory reasons can turn it on; everyone else gets the unmodified model. That split may become the standard operating model as different jurisdictions impose different rules — a balkanized AI output landscape where the same model behaves differently depending on where you access it.

There's a deeper tension here. Watermarking only works if detectors are widely available and trusted. OpenAI's decision to restrict detector access to approved institutions, while understandable from a security standpoint, limits the practical utility of the watermark for ordinary users and platforms. If social media companies, newsrooms, and educators can't run the detector themselves, the watermark exists in theory but doesn't change behavior in practice. The planned open-source release may resolve this, but until then, textGrain is a compliance checkbox more than a user-facing tool.

What to watch

The invisible watermark is a small technical change with a large signaling effect. AI text is about to become traceable — at least in Europe, at least for now.