AI

OpenAI Quietly Marks EU ChatGPT Text With a Hidden Statistical Fingerprint, and Light Editing Can Erase It

(today) · 4 min read · By Future Technology

Key takeaways

  • OpenAI's textGrain watermark biases word choices to leave a pattern that a detector can find later, and it will roll out to eligible EU ChatGPT and Codex output over the coming weeks.
  • OpenAI's own numbers show that swapping 25% of words for synonyms drops detection from about 92% to 17% on 400-token passages.
  • A missing watermark proves nothing about human authorship, and a detected one does not identify the user, account or prompt.

OpenAI is preparing to embed an invisible marker in the text that ChatGPT and Codex produce for users in the European Union. The company calls the system textGrain, and it is rolling out to eligible EU output over the coming weeks. The company has also been unusually direct about the limits of the approach, publishing figures that show how easily the mark can be weakened.

How textGrain works

Unlike a visible label or a hidden character tucked into the text, textGrain does not add anything you could spot or strip out by copying and pasting. Instead, it nudges the model's word choices as it writes. Language models pick each next word from a range of plausible options. A watermarking scheme tilts those choices slightly in a pattern that looks natural to a reader but is statistically unusual over a long enough passage. A detector that knows the pattern can then test whether a piece of text fits it better than chance would predict.

That design has one important consequence: the watermark lives in the overall distribution of words, not in any single phrase. It needs enough text to show up, and it can be diluted when the text changes.

OpenAI is not turning this on worldwide. Developers using the API can opt in for supported models starting now, but the feature is off by default. The company is also taking applications for access to its detector, though for the moment that tool will be limited to approved researchers and expert organizations. There is no public "paste text here and find out" checker, at least for now.

What the numbers say about reliability

OpenAI's published evaluation is a useful reality check for anyone hoping watermarks will settle the question of who wrote what. On 400-token passages (roughly 300 words), replacing 10% of the words with synonyms cut detection from about 92% to 66%. Replacing 25% of the words pushed it down to 17%. That is the kind of rewriting a student, marketer or content farm could do with a thesaurus tool or a second AI pass.

Length matters too. At a 1% false-positive target, meaning the detector wrongly flags human writing only one time in a hundred, OpenAI caught the watermark in about 80% of 200-token psychology answers. That rose to roughly 95% at 400 tokens. Short outputs such as a quick email reply or a code snippet give the detector far less to work with.

Subject matter also affects results. In areas like mathematics, where there are few valid ways to phrase an answer, the model has little room to vary its word choices, so there is little room to hide a pattern. Detection was noticeably worse there.

OpenAI states plainly that "the absence of a detected watermark does not prove human authorship." Text may be too short, edited or translated for the detector to work. The company adds that a positive result does not reveal who generated the text, which account or prompt was used, or how much a human contributed afterward. In other words, the tool can suggest that OpenAI's models were involved, but it cannot assign blame or measure effort.

On quality, OpenAI says benchmark results for GPT-6 Astra stay broadly similar with textGrain enabled, so users should not expect worse answers in exchange for the marker.

Why the EU, and what comes next

The company has not framed the EU launch in detail in the material available, but the geography is telling. European regulators have been pushing for greater transparency around AI-generated content, and a technical mechanism for marking machine-made text fits that direction. Starting in the EU lets OpenAI meet that pressure in the region where it is most acute while keeping the global default unchanged and watching how the system performs.

The practical value of the scheme is likely to be narrow but real. Researchers and trust-and-safety teams could use a detector to estimate how much unedited model output appears in a large body of text, such as a flood of spam, coordinated posts or bulk submissions. Statistical tests work better across thousands of documents than for judging a single essay. That is a very different job from policing individual authors, and the false-positive caveat should make anyone wary of using a detector result as proof against one person.

The bigger lesson is that watermarking is one layer, not a verdict machine. Anyone willing to paraphrase, translate or heavily edit output can weaken it, and text from other vendors or open models will not carry OpenAI's mark at all. Schools, employers and publishers that hope for a clean yes-or-no answer will not get one from textGrain. What they may get is another signal, to be weighed alongside context, drafts and old-fashioned conversation.

Sources

More from Future Technology