OpenAI adds invisible watermarks to ChatGPT and Codex text in the EU
OpenAI will watermark eligible ChatGPT and Codex text for EU users in the coming weeks and lets API customers worldwide opt in. Its own tests show the signal fades in short or edited text.

If you use ChatGPT or Codex in the European Union, the text you get back is about to carry a hidden signature. On Monday, 5 October 2026, OpenAI explained how it will meet the text provenance rules of the EU AI Act. Over the coming weeks it will add an invisible watermark to eligible ChatGPT and Codex output for users in the EU, on every plan. At the same time, API customers anywhere in the world can now switch on the same watermark for selected models, and researchers can apply for access to a detector that looks for it.
The technology is called textGrain. You will not see it, you cannot strip it out by deleting odd characters, and it does not change what the model says in any obvious way. What it changes is which words the model picks, in a pattern that only a detector with a secret key can recognise. OpenAI is also unusually open about how weak the signal can be, and that honesty is the most interesting part of the announcement.
What OpenAI announced
The announcement came as a post on the OpenAI blog titled "Our approach to EU text provenance rules", together with a help center article and a technical report. It has three parts.
- EU rollout in ChatGPT and Codex. Over the coming weeks, eligible text output in ChatGPT and Codex gets a watermark for users in the EU, across all plans. OpenAI says it is not making this a global default at launch, because a regional start lets it learn from real use and feedback.
- Opt in for API customers worldwide. Starting on 5 October, developers who use the OpenAI API can turn on watermarked output for selected models. It stays off by default. OpenAI says it is also working with cloud partners so that OpenAI models accessed through their services can be watermarked in the coming weeks.
- A detector for a small group. Approved researchers and expert organisations can now apply for access to the textGrain detector. Access is granted case by case, in line with the EU Code of Practice on transparency of AI generated content. The tool only reports whether it finds an OpenAI watermark. It does not identify the user and does not reveal prompts or conversations.
OpenAI stresses that all of this is about text only. Its tools for images and audio, including the openai.com/verify web page and the Content Provenance API, stay publicly available.

Why the EU is forcing this
The trigger is Article 50 of the EU AI Act. Its transparency rules have applied since 2 August 2026. Paragraph 2 says that providers of generative AI systems must make their output machine readable and detectable as artificially generated or manipulated, using technical solutions that are effective, interoperable, robust and reliable as far as that is technically feasible. That last phrase matters a lot, as we will see below.
To show how providers can meet the rule in practice, independent experts wrote a Code of Practice on transparency of AI generated content, in a process run by the EU AI Office. The final version was published on 10 June 2026. Signing it is voluntary, but the obligations in Article 50 are not. According to the European Commission, about 190 companies and organisations had signed the code by the end of July 2026, among them OpenAI, Anthropic, Google, Meta, Microsoft and Mistral.
In other words, OpenAI did not wake up one morning and decide that text watermarks are a great idea. It had to deliver something, and Monday's post is its answer.
How textGrain works
To understand the trick, it helps to remember how a language model writes. It does not pick whole sentences. It produces one token at a time, a word or a piece of a word, and for each step it has a list of possible next tokens with probabilities. Normally a bit of randomness decides which one is chosen, so the same prompt can give slightly different answers.
A statistical watermark replaces that ordinary randomness with randomness that comes from a secret key. The model still prefers the same likely words, but whenever several choices are almost equally good, the key quietly nudges the pick in a direction that a detector can later check. Read one sentence and nothing looks strange. Read a few hundred tokens with the key in hand and the pattern becomes measurable.
OpenAI's technical report, written with researchers from the University of Pennsylvania and Yale University, gives the maths. It describes the method as coupling token generation to keyed randomness through an optimal transport problem, with costs based on Gumbel random variables and a regularisation term based on the Kullback Leibler divergence. The practical point is simpler: the watermark has a strength setting, called an entropy budget, that says how much of the model's natural randomness is given up in exchange for the signal. A stronger watermark is easier to detect but takes away more freedom from the model. To keep the computation fast, the method groups the vocabulary into blocks and keeps the relative probabilities of tokens inside each block unchanged.
The detector needs only two things: the text and the secret key. It does not need the original model or the strength setting that was used. OpenAI says it will publish textGrain as open source so others can build on it, and that the technical report will be updated with more details in the coming weeks.
How well it works, by OpenAI's own numbers
This is where the announcement gets refreshingly frank. OpenAI says textGrain matched or beat the other methods it tested, including Google DeepMind's SynthID for text. Then it immediately adds that strong performance under ideal conditions does not guarantee reliable detection in daily use, and it publishes numbers that show exactly why.
All figures below come from OpenAI's tests on English answers to questions from the ELI5 dataset, at a target false positive rate of 1 percent. That means the detector is tuned so that about one in a hundred texts without a watermark would wrongly be flagged.
- Length matters. For topics such as psychology, the detector found the watermark in about 80 percent of 200 token passages and about 95 percent of 400 token passages.
- Rigid topics are harder. For mathematics, where there is much less freedom in word choice, detection was substantially lower. OpenAI does not give an exact number in the blog post.
- Editing hurts a lot. In 400 token passages, replacing 10 percent of the words with synonyms cut detection from about 92 percent to 66 percent. Replacing 25 percent of the words cut it to 17 percent.

A token is roughly three quarters of an English word, so 400 tokens is about 300 words. A short email reply, a chat message or a few lines of code are well below that, which is why short outputs are the hardest case. A sentence or two simply does not contain enough choices for the pattern to show up with confidence.
OpenAI also checked whether the watermark makes the model worse. Across the benchmarks it uses for Astra, its latest frontier model, it reports no meaningful difference. On the Artificial Analysis Intelligence Index, Astra scored 49.57 points without the watermark and 49.76 with it. On GPQA Diamond, the scores were 94.44 percent and 93.94 percent. On Terminal Bench 4.0 the watermarked version even scored higher, 56.06 percent against 53.90 percent. These small ups and downs look like normal noise, which is the point OpenAI wants to make.
The Register, which called the feature "weak sauce" in its headline, raised one question the report does not answer: whether a nudged word choice could ever change the meaning of a sentence, for example when the alternative word carries a different fact or number. Anthropic's method nudges word choices too, so the same question applies there.
What a watermark cannot tell you
OpenAI spends a whole section on the limits, and it is worth repeating them in plain words, because these are exactly the questions teachers, editors and employers will ask.
- A watermark does not measure how much a human contributed. It can suggest that an OpenAI system generated or processed part of a text, but not how much thinking, editing or creativity a person added.
- It does not settle ownership or responsibility. It cannot say who owns the text, whether its use was legal, or whether it had to be disclosed.
- It does not identify anyone. No person, account, prompt or conversation is linked to the text.
- It does not check facts. A watermarked text can be true or false, and so can an unmarked one.
- No watermark does not mean a human wrote it. The text may be too short, edited, translated, made with an older or unsupported model, or written with another company's tool.
Because missed watermarks and false alarms are both possible, OpenAI is not releasing the detector to the public at launch. That is a sensible call. A public "AI or not" checker with these error rates would quickly be used to accuse students and writers of cheating, and some of those accusations would be wrong.
How this compares to Anthropic and Google
OpenAI is not the first to do this, and the comparison shows how differently the big labs read the same law.

Anthropic announced its approach in August. Claude's text watermark is a version of SynthID Text, the method Google DeepMind published in Nature in 2024. Anthropic applies it worldwide, not just in Europe, because it says it does not yet have a durable way to limit the watermark to one region. Claude models launched on or after 2 August 2026 carry the mark from day one, and Anthropic is adding it to older models over the coming months. Its detector runs as a private preview for groups such as regulators, fact checkers, researchers and companies with their own compliance duties.
Google itself has used SynthID Text in Gemini since 2024 and also marks images and audio with SynthID. OpenAI already uses SynthID watermarks and Content Credentials for supported images and audio, so it is not new to provenance. For text, though, it built its own method and chose a much narrower rollout: automatic only in the EU, optional everywhere else.
That difference is the real story. A user in Berlin will get watermarked ChatGPT text, a user in New York will not, unless a developer has switched it on in the API. With Claude, both users get the same watermark.
Why it matters
This matters for three groups in particular.
First, for schools and universities in Europe. Many teachers hope for a reliable way to spot AI written homework. This is not it, at least not yet. The detector is not public, short answers are hard to check, and a student who rewrites a quarter of the words wipes out most of the signal. Anyone who treats a watermark result as proof is making a mistake that OpenAI itself warns against.
Second, for businesses and creators. If you publish content drafted with ChatGPT in the EU, it may now carry a signal that a detector can find later. That does not make it illegal or bad, and the watermark says nothing about quality or ownership. But companies that care about how their content is labelled should know that the marking is happening, and developers who build products on the API now have to make an active choice about whether to turn it on.
Third, for the AI Act itself. Article 50 asks for marking that is robust and reliable as far as technically feasible. OpenAI's own numbers show that robust is a stretch once text is edited. Regulators will have to decide whether this is good enough or whether the bar will rise over time. OpenAI clearly expects that, and says it will revisit every part of the approach as the technology, standards and evidence develop.

Dany's take
I write about AI every day, and I use AI tools to do it, so this one hits close to home. I think watermarks are a fair idea in principle. People should be able to find out when a machine wrote something, especially with elections, fake news and deepfakes around. What I like about OpenAI's post is the honesty. They openly say the signal is weak in short texts and breaks down when someone edits the words. That is far better than marketing a magic AI detector that ends up getting innocent students in trouble. What I find strange is the split between Europe and the rest of the world. As someone who lives in Germany, I now get watermarked ChatGPT text while users in the US do not. Anthropic just marks everything everywhere, which feels more consistent. My practical advice: do not panic, and do not trust any tool that claims to prove a text was written by AI. If you use ChatGPT for drafts, edit them and make them your own. That was good advice before the watermark, and it still is.
Sources: OpenAI: Our approach to EU text provenance rules and the textGrain technical report. Also: The Verge: OpenAI is adding text watermarking in ChatGPT and Codex, The Register: OpenAI rolls out weak sauce watermarking for AI text, Unite.AI: OpenAI begins phased text watermarking under EU AI Act rules and Anthropic: How Claude's text watermarking works.
Source: openai.com