Change one word in ten, and OpenAI’s text watermark falls from 92% to 66%

2026-10-07

On 5 October 2026 OpenAI turned on a text watermark (a hidden pattern that lets a matching detector guess the text came from their model) called textGrain. It is not a new chatbot. It is a mark inside the words a chatbot already writes.

Two doors opened. Worldwide, API customers (people who pay to send text to the model from their own software) can opt in for select models. The mark stays off unless they switch it on. Over the coming weeks, eligible ChatGPT and Codex text for people in the European Union gets the mark automatically. The detector (the checker that looks for the mark) is not a public website. OpenAI is taking applications from approved researchers and expert organizations first.

OpenAI textGrain rollout: API opt-in worldwide, automatic for eligible ChatGPT and Codex text in the EU, detector not public

The number that should change how you read a headline is this one, from OpenAI’s own post, not from a lab that is selling you a detector. On 400-token passages (a token is a small chunk of text, often part of a word; 400 tokens is on the order of 300 English words, not a fixed page), swapping 10 percent of the words for synonyms cut detection from about 92 percent to 66 percent. Swapping 25 percent cut it to 17 percent. A missing mark does not prove a human wrote the page. OpenAI says that sentence itself.

Start where the model starts: one next piece of text

Simple mindmap of OpenAI textGrain: word-choice watermark, secret key, closed detector, 92 to 66 percent after 10 percent synonyms, EU Article 50, a miss is not proof

A large language model (a program that continues text by guessing the next chunk, over and over) does not look up a finished essay. It holds a list of possible next tokens and a probability (a weight, like how many slips of paper each word gets in a hat) for each one. Then it draws.

After the words “The morning was”, many next words are fine: warm, cold, mild, calm, sunny, bright. None of those is a hidden character. They are ordinary words a reader expects. That pile of acceptable next words is the only place a text watermark can hide, because the page you copy is just words.

A language model picks the next word from a hat of acceptable words such as warm, cold and mild

If the next word is forced, the hat has one slip. Two plus two is four. A line of code that must be return has few honest substitutes. A psychology explanation has many. The watermark spends freedom. Where there is almost no freedom, there is almost no place to hide a pattern. Hold that. It is why the math numbers later look worse than the essay numbers.

This is not invisible ink

Older tricks stuffed a secret into the file: a white letter on a white background, a weird space, a character your eye skips. Copy the words into a new document by hand and that trick dies. textGrain does not add hidden characters, extra spaces, or odd punctuation. OpenAI’s description is a statistical signal (a tilt in which ordinary word gets picked, visible only if you know the secret and you count) in the model’s word choices.

Paste still carries the mark, because the words themselves are the mark. A rewrite can strip it, because a rewrite picks different words. That is the whole product, stated before any law.

Invisible ink tricks die on retyping, while the textGrain word-choice mark survives paste and weakens with edits

How the secret actually steers one word

You need three ideas, in order.

A secret key (a password the writer and the detector share, and you do not) plus the last few words decide a pattern for this position. The technical report of 5 October 2026, by Xiang Li, Garrett Wen, Xiaohong Chen, Qi Long, Arzav Jain, Florent Joly, Mike Lam, Qingquan Song, and Weijie Su, says that pattern splits the vocabulary (every token the model is allowed to emit) into blocks. The key picks a favored block. Inside that block, tokens keep their original relative odds. The model still says something sensible. It is nudged toward the block the key wanted at that spot.

The report’s figure uses ordinary next words such as warm, cold, calm, and sunny as the blocks. You do not need their exact percentages to see the trick. “Cold” and “mild” can both be honest. The key makes one of those honest choices a little more likely, on purpose, in a way that flips as the sentence moves.

textGrain secret key and recent words split the vocabulary into blocks and favor one block

Unbiased (if you averaged every possible key, the hat would look like the original model again) is why they can claim the writing did not get worse. The tilt depends on a key you do not hold. Over all keys, the tilts cancel. One essay still carries one key’s tilt. That is the signal.

The detector does not need the model, and the paper says it does not need the entropy budget used while writing. It needs the text and the same secret key. It rebuilds the favored blocks and asks a counting question: did the words land in the favored blocks more often than luck? If yes, it reports an OpenAI watermark. It does not report who typed the prompt, which chat it came from, or whether the sentence is true.

textGrain detector uses the passage and the same secret key to count favored words against chance

The budget: how much freedom the mark is allowed to spend

Entropy (here, a measure of how many different next words are still reasonable) is the raw material. A weather sentence has high entropy. A formula has low entropy. The paper sets a budget, a fraction of that freedom the watermark may spend, written as the Greek letter beta. Spend none, and there is no pattern to find. Spend all of it, and the sentence starts sounding like the key, not like the model.

OpenAI says textGrain matched or beat the other methods they tried, including SynthID for text (Google’s text watermark, a different recipe with the same job). They also say a good score in a clean test is not a promise about a messy document. Believe both lines. The second one is the one a school or a newsroom will wish they had read.

High entropy essays leave room for a watermark, low entropy math and code leave little room

The only detection table that counts

These figures are OpenAI’s, on English answers, with the detector aimed at a 1 percent false positive rate (out of 100 texts that were never watermarked, the checker is tuned to cry wolf about once). They are not an independent rerun. Psychology here means the flexible kind of ELI5 answer (Explain Like I’m Five, a set of plain-language questions), not a clinical claim about the field.

SituationDetection, as OpenAI states itHow to read it
About 400 tokens, psychology-styleabout 95 percentA clean, long, flexible answer. The best case they printed.
About 200 tokens, same kind of textabout 80 percentHalf the length, and one in five marks is already missed.
Mathematicssubstantially lower. No single percent in the prose next to those two linesThe hat is smaller. Do not invent a number they did not print.
400 tokens, 10 percent of words swapped for synonymsabout 92 percent down to 66 percentA light edit. Still their test, not a clever attack.
400 tokens, 25 percent of words swapped17 percentMost marks are gone. The page can still read like the same answer.
OpenAI textGrain detection falls from about 95 percent to 66 percent with 10 percent synonyms and 17 percent with 25 percent

The New Stack quotes OpenAI that code is harder for the same reason as math: fewer plausible next choices. I am not going to pretend their coding-detection rate was printed beside the 95 percent. It was not, in the sentences above. Codex in the EU is still scheduled to get the mark. A mark that is hard to find later is a weak receipt.

A false positive rate of 1 percent is not a vibe. Point the checker at 100 unwatermarked passages, expect about one alarm. In a school of a thousand essays, that tuning alone is about ten false alarms before anyone has cheated. OpenAI is not handing you this checker yet. The rate is still the right size of doubt for the day they do.

Did the writing get worse? Their table says no, and it is their table

OpenAI compared Astra at maximum with the watermark off and on. They call the gaps not meaningful. Read the pairs before you repeat the slogan. Some watermarked scores are a hair higher. That does not mean the mark makes the model smarter. It means this test did not show a quality tax.

TestAstra, no watermarkAstra, watermarked
Artificial Analysis Intelligence Index49.5749.76
AutomationBench34.09 percent34.86 percent
DeepSWE v1.172.80 percent71.68 percent
Terminal-Bench 4.053.90 percent56.06 percent
Terminal-Bench Science 0.156.90 percent60.00 percent
BrowseComp87.92 percent87.35 percent
HealthBench Professional64.27 percent64.60 percent
GPQA Diamond94.44 percent93.94 percent
Watermark on versus off: small quality gaps in both directions, so no quality tax on this table

The unbiased property is the reason a flat table is even plausible. The key spends a little freedom at each word and, across keys, gives it back. Your one document is not the average. It is one key’s path. Quality can stay flat while the path is still detectable. Those are different measurements.

What a missing mark is allowed to mean

OpenAI’s sentence, kept whole: the absence of a detected watermark does not prove human authorship. The text may be too short, edited, or translated. It may come from a model they do not mark, from before this switch, or from another company.

A hit is also not a name. The tool, as OpenAI describes it, says whether an OpenAI watermark is present. It does not say which account, which prompt, or how much a person edited afterward. “A watermark does not measure human contribution” is their line. A student who pasted a paragraph into a larger essay, and a model that wrote the entire essay, can both produce a hit. A hit is a lead. It is not a fraction of authorship.

What a missing or detected textGrain watermark does and does not mean

There is no public number in the 5 October post for translation. They list it as a reason detection fails. Do not invent a percent. The synonym test is the one they quantified: words changed inside the same language already do most of the damage.

The law that made them ship it, in one page

Article 50 of the EU AI Act (the Union’s law on artificial intelligence; this article is the transparency chapter, not the high-risk chapter) applies from 2 August 2026. The European Commission’s fact page says providers must put a machine-readable mark (a mark a program can check, not only a sentence a person can read) on synthetic text, image, video, and audio, and must make that mark detectable. An assistive edit that does not really change the input is outside that duty. Systems already on the market before 2 August 2026 get a grace period for the marking duty until December 2026. Legal summaries of the amendment put the last day at 2 December 2026. The Commission page itself says December 2026. I will not pretend those two sentences are the same document.

ChatGPT telling you it is a bot is a different duty from marking the paragraph it writes. A person can see the first. Only a detector with the key can see the second. The EU also publishes visible icons for AI-made content. textGrain is not that icon. One is for eyes. One is for a counter.

A visible AI label for people compared with a machine-readable watermark only a key holder can check

What you should do with this, depending on your job

Do not build a cheating policy, a newsroom rule, or an API default on the 95 percent cell. That cell is a long, clean, flexible English answer that nobody edited, checked by the company that made the mark.

Your jobUse the mark forDo not use it for
Teacher or integrity officeA lead, once a real detector exists for you, on a long unedited passageProof. A miss is not innocence. A hit is not a named student or a percent of the essay.
StudentKnowing the mark is not a moral fact. It is a fragile statisticA plan. Swapping synonyms to dodge a checker is still the same assignment problem. This piece is not a recipe.
API developerAn opt-in switch, off by default, if a customer in scope asks for a receiptAssuming opt-in gives you the detector. OpenAI says it does not.
Newsroom or publisherOne check among others, on text that is still close to the model outputA published claim that unmarked text is human. Other companies, edits, and translations all produce misses.
Anyone outside the EU using ChatGPT in the browserNothing, yet, unless you are in the rolloutAssuming your chat is already marked. The automatic switch described here is the EU one.
Next steps after the watermark detector fires or stays quiet

If you run the API and you do nothing, you are on the default: no textGrain. Turning it on does not make you the detective. Detection stays with OpenAI’s approved-access list while they decide the false alarms are understood. That split is the part vendors will blur in a sales call. Write it into the contract before you promise a customer a receipt you cannot read.

What is still secret, so do not fill it in

OpenAI says the report will be updated in the coming weeks, and that they plan to release the technology as open source (the method published so others can rebuild it). There is no date on that release. There is no public list, in the post, of which API models are eligible. There is no printed detection rate for code, for translation, or for languages other than the English evaluation they described. There is no price for the detector. There is no switch a normal ChatGPT user can flip to show the mark on their own paragraph.

What is known and still unknown about OpenAI textGrain: code rate, translation, languages, models, open-source date

The useful object is not the name textGrain. It is the shape of the failure. The mark lives in spare word choice. Long, flexible, unedited English is where their own detector looks strong. Short text, math, code, and a rewrite are where it goes quiet. Quiet is not a person.

Common questions about OpenAI textGrain

What is OpenAI textGrain?

It is a text watermark OpenAI turned on on 5 October 2026. It adds no hidden characters. A secret key nudges which ordinary word the model picks, and a detector with the same key counts whether favored words show up more often than chance.

Can I check a passage for the textGrain watermark?

Not yet. The detector is not public. OpenAI is taking applications from approved researchers and expert organizations first, and API customers who opt in do not get the detector.

Does editing remove the watermark?

Mostly, yes. In OpenAI’s own test on 400-token passages, swapping 10 percent of the words for synonyms cut detection from about 92 percent to 66 percent, and swapping 25 percent cut it to 17 percent.

Does no watermark mean a human wrote it?

No. OpenAI says a missing watermark does not prove human authorship. The text may be too short, edited, or translated, or it may come from another company, an unmarked model, or from before the switch.

Leave a comment