Breaking
AI & MLConfirmed

OpenAI's EU Watermark Push Has Limits

OpenAI is watermarking ChatGPT and Codex text in the EU, but its own tests show editing quickly breaks detection.

··2 hours ago·4 min read
a computer generated image of a human head
Photo by Growtika on Unsplash

OpenAI is preparing to mark AI-generated text in the European Union with a signal most readers will never see. According to the company, text produced by ChatGPT and Codex will soon carry an invisible watermark — a statistical pattern woven into word choices rather than any visible label on the page. The catch, per OpenAI's own testing, is that the mark tends to fall apart under ordinary editing.

How the watermark actually works

OpenAI describes the mechanism as a technology it calls textGrain. Rather than inserting hidden characters or metadata, the system nudges the model's word choices to create a detectable statistical pattern. The watermark does not change how the text reads or copies, and it is not visible to anyone looking at the output.

Detection works by looking for that pattern. OpenAI is not saying the mark identifies a specific account, prompt, or conversation — only that the text shares the statistical signature. That distinction matters for anyone expecting forensic certainty about who wrote what.

The EU rollout and its timing

OpenAI says the EU deployment will roll out gradually. "Over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union," the company explained. Eligible output, not all output, is the operative phrase; OpenAI has not said the watermark applies to every response.

The company is not making watermarking a global default. API developers worldwide can opt in for supported models, but the feature remains disabled by default outside the EU plan.

Editing weakens the signal fast

OpenAI's own evaluation data shows how brittle the approach can be. In tests using 400-token passages, replacing 10% of words with synonyms cut detection from about 92% to 66%. Replacing 25% of the words dropped it to 17%. Those are edits of the kind a person might make while polishing a draft.

The company also measured how length affects reliability. At a 1% false-positive target, OpenAI detected the watermark in about 80% of 200-token psychology responses, compared with roughly 95% when the text reached 400 tokens. OpenAI published a chart showing results for watermarked responses to mathematics and psychology questions from the ELI5 dataset at that same 1% false-positive rate.

Subjects where the model has less flexibility in wording made detection harder still. Mathematics was worse than psychology in the company's tests, because the model has fewer alternative ways to phrase the same answer.

What a missing mark does not prove

OpenAI is explicit that the absence of a watermark is not evidence of human authorship. "The absence of a detected watermark does not prove human authorship," the company warned. "Text generated with OpenAI tools may be too short, edited, or translated for detection to work reliably."

That caveat cuts against the most obvious use case for watermarking — deciding whether a given piece of writing came from a machine. A short response, a translated passage, or a heavily edited draft can all carry the mark and still evade detection, according to OpenAI's own framing.

The reverse is also limited. OpenAI says a detected watermark does not reveal who generated the text, their account, their prompt, or the conversation behind it. Nor can it say how much of the final work a human wrote or edited. The mark speaks to provenance at the text level, not authorship at the person level.

Quality held steady in benchmarks

One concern with any watermarking scheme is that it degrades the output. OpenAI says that is not the case here. Watermarking does not meaningfully affect the quality of GPT-6 Astra, with benchmark results remaining broadly similar when textGrain is enabled.

That claim rests on the company's internal benchmarks. Independent verification of watermark robustness and quality impact is not part of the announcement.

Who gets to run the detector

OpenAI is also opening applications for its watermark detector, but access will initially be limited to approved researchers and expert organizations. There is no general tool for the public to check whether a document was AI-generated.

The gatekeeping means the detector will be used in research and expert settings rather than by teachers, editors, or employers trying to vet text at scale.

What the numbers say

The company's published figures sketch the tradeoffs:

  • Replacing 10% of words with synonyms in 400-token passages cut detection from about 92% to 66%.
  • Replacing 25% of words dropped detection to 17%.
  • At a 1% false-positive target, detection was about 80% for 200-token psychology responses, versus roughly 95% at 400 tokens.

Those numbers describe an evaluation of watermarked responses, not a guarantee about any individual piece of text. OpenAI's framing throughout is probabilistic: the mark can be detected under some conditions, and it can fail under others.

What this means for readers and businesses

The EU rollout will put a detection signal into ChatGPT and Codex output for one region first, with API access opt-in everywhere and off by default. For organizations that were hoping watermarking would settle questions of authorship, OpenAI's own caveats suggest it will not. A clean detector result may still be inconclusive, since short, edited, or translated text can evade the mark, and a detected mark does not name a person or account.

That leaves the practical burden where it already sits: on editors, researchers, and reviewers who have to weigh the text itself rather than rely on a single automated verdict. The company's choice to limit detector access to approved researchers and expert organizations reinforces that the tool is meant for careful evaluation, not routine gatekeeping. For now, the watermark is one signal among several — and, per OpenAI's own testing, a fragile one.

Reporting based on original coverage from BleepingComputer.

#openai#watermarking#chatgpt#eu#ai detection

Sources

Iliyas

Founder & Editor, Xploitwire

This article was written and reviewed against the sources listed above before publication, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories