Yes, hidden characters turn up in text produced by AI assistants, and almost nothing else about this question has been measured by anyone ranking for it. No OpenAI model was sampled for this page, so it does not claim what ChatGPT’s output contains, and that puts it ahead of the search results rather than behind them: not one of the eight pages ranking for this question publishes a corpus size, a sampling method or a count either. What can be measured has been. Across the seven of those pages that could be extracted, exactly one character is named by a majority. Across 436 web pages and 809,229 words, zero-width characters appear on 7.8% of pages and the curly apostrophe on 55%, with 56 of those on a blog post from 2017. And Google’s own AI Overview took both sides of the watermark question within minutes.
Which characters do the ranking pages name?
One character is named by a majority of the pages ranking for this query: U+200B ZERO WIDTH SPACE, at five of seven. Nothing else reaches four and a half. The count below comes from a regular-expression pass over the seven extracted pages, searching for each character by its codepoint number and by its common name, counting pages rather than mentions.
| Character | Pages naming it, of 7 |
|---|---|
| U+200B ZERO WIDTH SPACE | 5 |
| U+200C ZERO WIDTH NON-JOINER | 4 |
| U+200D ZERO WIDTH JOINER | 4 |
| U+FEFF byte order mark | 4 |
| U+00AD SOFT HYPHEN | 4 |
| U+00A0 NO-BREAK SPACE | 3 |
| U+2014 EM DASH | 3 |
| Curly quotes, U+2019 and U+201C | 2 |
| U+202F NARROW NO-BREAK SPACE | 2 |
| U+2060 WORD JOINER | 2 |
Two readings matter more than the individual rows. The first is that a category with ten candidate characters and one majority answer has not converged on what it is describing. The second is sharper: two of the seven pages, medium.com at position four and itc.ua at position seven, name no codepoint at all while ranking for a query that asks which characters are added.

What is each of these characters actually for?
Every character on that list has a job, and three of them have jobs in living writing systems. No page in the ranking set says so, and it changes what finding one of them means.
- U+200D ZERO WIDTH JOINER fuses separate emoji into one. A family emoji or a profession emoji is several people or objects joined by this character, which is why it appears in ordinary messages constantly.
- U+200C ZERO WIDTH NON-JOINER is required to write Persian and Hindi correctly, where it stops two letters from joining into a single connected form. A detector that flags it flags Persian.
- U+00AD SOFT HYPHEN marks a place where a word may break across lines. It prints only when the break happens, which is exactly why it is invisible.
- U+FEFF at the start of a file is a byte order mark, which labels the encoding. Editors and export tools add it without asking.
- U+00A0 and U+202F, the non-breaking spaces, hold a number and its unit on the same line, so a price or a measurement does not split across a line break.
- U+200B ZERO WIDTH SPACE marks a permitted break with no hyphen, and it is also the one character in the list with no everyday typographic use, which is why it is the one most worth looking at.
The consequence is a rule that survives whatever anyone decides about watermarks. A character that performs a linguistic function in a living script is not evidence about who typed the sentence, and the invisible character detector names each one it finds rather than scoring the text.
How common are these characters in ordinary web pages?
Across 436 competitor pages extracted from live search results into this project between 3 and 23 August 2026, totalling 809,229 words, at least one zero-width or soft-hyphen character appears on 34 pages, which is 7.8%. The corpus is every page pulled whole into this project’s research folders during that period, excluding the analytical summaries written here, with the extraction provenance comments stripped before counting.
| Character | Occurrences | Pages carrying it, of 436 |
|---|---|---|
| U+200B ZERO WIDTH SPACE | 83 | 9 |
| U+200C ZERO WIDTH NON-JOINER | 113 | 2 |
| U+200D ZERO WIDTH JOINER | 97 | 19 |
| U+FEFF | 2 | 2 |
| U+00AD SOFT HYPHEN | 178 | 5 |
| U+00A0 NO-BREAK SPACE | 35 | 5 |
| U+2060 WORD JOINER | 1 | 1 |
| U+2014 EM DASH | 1,813 | 172 |
| U+2013 EN DASH | 753 | 142 |
| U+2019 curly apostrophe | 3,776 | 240 |
| U+201C opening curly quote | 967 | 155 |
| U+201D closing curly quote | 1,254 | 155 |
The limit belongs here rather than in a footnote. These pages were collected from live search results in 2026 and their authorship is unknown, so this corpus cannot separate human writing from machine writing, and nothing above is offered as a measurement of human prose. What it establishes is a base rate: these characters are ordinary in web text, whoever produced it. A signal present on 7.8% of pages before anyone asks about AI is a weak signal about anything.

Is punctuation evidence of anything?
The two punctuation claims have different answers and are usually made in the same breath. The curly apostrophe is not an AI artifact and nothing supports treating it as one. It appears on 240 of the 436 pages, 55.0%, and 56 of those occurrences sit on a single blog post published in September 2017, five years before ChatGPT was released. Word processors have converted straight apostrophes to curly ones automatically for three decades.
The em dash is less settled, and this page says so rather than picking the convenient reading. It appears on 172 pages, 39.4%, in a corpus whose authorship is unknown. Three pages in that corpus carry visible publication dates predating ChatGPT and were extracted rather than summarised here, and across those three there are zero em dashes and 58 curly apostrophes. Three pages is not a sample that settles anything, in either direction, and treating it as one would repeat the mistake the rest of this SERP makes.
This site removes em dashes from its own writing, for a reason that has nothing to do with detection: a comma, a colon or a full stop states the relationship between two clauses, and an em dash leaves it to the reader. The em dash remover does it mechanically for anyone who wants the same.
Watermark or artifact: what can actually be said?
Does ChatGPT add hidden characters to text on purpose, or incidentally? The ranking set is split, Google’s AI Overview took both sides within minutes, and nothing on the search results page settles it. Naming the split: medium.com, the LinkedIn post at position two, humanwritesai.com and wipe-ai.com describe an intentional watermark, and three of those four sell a tool that removes it. The top answer in the Reddit thread at position one says the characters “are not inserted randomly or to covertly express a certain pattern”, and Google’s own first summary agrees with the Reddit side.
The instability is worth stating exactly, because it is the reason this page attributes nothing to OpenAI. On the first capture, Google’s AI Overview stated that OpenAI has said these artifacts are accidental side effects of training via reinforcement learning rather than an intentional tracking watermark. On a second capture minutes later, from the same query with the same parameters, that attribution was gone, replaced by a line about Reddit users debating whether the characters are watermarks or tokenizer quirks. A statement attributed to a named company on one run and absent from the next is not a source, and repeating it would be publishing a coin flip as a fact.
What can be said is narrower and more useful. Hidden characters do appear in AI output. They also appear in web pages, word processors, translation memories and every script that needs them. Nobody ranking for this question has published a rate for either population, so the question of whether a zero-width space in your document came from a model is currently unanswered by anyone, including this page.

What one model’s output actually contained
90,880 words of Claude output, across 28 full-length drafts plus a passage generated straight to a file, were scanned on 18 August 2026 for zero-width characters, unusual spaces, bidi controls, variation selectors and tag-block characters, and returned zero. Claude is not ChatGPT. That measurement was run while building the Claude watermark remover and it says nothing about any OpenAI model.
It is here for two reasons. It is the only corpus-sized measurement of any model’s raw output anywhere in this research, which makes the absence of one on eight ranking pages easier to see. And it carries the mechanism that the whole category confuses: a statistical watermark lives in which words the model chose, not in anything hidden between them, so a tool that strips invisible characters cannot remove it and a scan for invisible characters cannot find it. Those are two different objects with one name.
How do you check a document for them?
Paste the text into a detector that lists codepoints, because these characters are invisible by definition and no amount of careful reading finds them. That is the whole procedure, and the invisible character detector reports what is present and what each one is for.
The constraint on the result is the part nobody states. What comes back is a list of what is in your document, which tells you nothing about where it came from, because the same characters arrive from web pages, word processors, translation tools, emoji sequences and every language that requires them. A detector answers what, never who.
Can you remove them, and does removing them do anything?
Removing them is easy and worth doing whenever the characters break something concrete: a search that will not match a name, a CSV column that will not parse, a form field that rejects input that looks correct, a filename that behaves oddly. Those are real problems with a mechanical fix, and the AI text cleaner applies it.
Removing them does not make text undetectable, and this page declines that claim as the three AI tool pages on this site have. Nothing in this research supports it, and the mechanism argues against it: if a model’s watermark is statistical, it survives every character you strip, and if the characters were incidental to begin with then removing them changes nothing about the text’s origin. The AI text detector here returns no score for the same reason, which is that a classifier without a published accuracy figure is a guess with a number on it.
What does each removal method actually remove?
The three methods Google’s AI Overview recommends do three different things, and only one of them touches codepoints at all. Pasting as plain text strips rich-text formatting, bold, links and fonts, and keeps every character in the string, invisible ones included. A plain text editor does exactly the same: it has no formatting to strip, and it stores the characters it is given. Ctrl+Shift+8 in Microsoft Word displays formatting marks, which shows paragraph marks, tabs and spaces rather than revealing arbitrary invisible Unicode.
So two of the three recommendations leave the characters exactly where they were, and the third shows some of them at best. What works is a tool that operates on codepoints, and if the goal is stripping markdown syntax rather than invisible characters, the markdown stripper is the one that does that job.
Where does each of these questions get answered?
Each page below answers one question in this section, so the list routes by question rather than by title.
- What an invisible character is and what it breaks, where none of the ten pages ranking for that query measures anything.
- Which invisible characters a document contains, which names each one rather than scoring the text.
- How to strip them out, for when a character is breaking something concrete.
- How to remove markdown syntax, which is a different job from removing characters.
- How to convert curly quotes back to straight ones, the punctuation half of the same cleanup.
- Why this site’s detector returns no score, where four ranking detectors publish false positive rates spanning 0.01% to 1% and only one attaches a figure to a named population.
Questions people ask about this
These are the five questions Google served in its People Also Ask box across two captures of this query on 23 August 2026. Four are answered below; the fifth is out of scope and says why rather than being quietly dropped.
Does AI text have invisible characters?
Sometimes, and so does ordinary web text. Across 436 pages and 809,229 words measured here, 7.8% carry at least one zero-width or soft-hyphen character, and the authorship of those pages is unknown. A scan of 90,880 words of Claude output returned none at all. Finding one in a document tells you the document contains it, not where the document came from.
How do I remove a ChatGPT invisible watermark?
You can remove invisible characters from any text in seconds, and that is not the same as removing a watermark. A statistical watermark is carried by which words the model chose, so it survives every character you strip. Removing invisible characters is worth doing when they break a search, a parser or a form field, and it is not a route to undetectable text.
How can I remove Unicode characters from ChatGPT text?
Use a tool that operates on codepoints rather than on formatting. Pasting as plain text and retyping in a plain text editor both keep every character in the string, which is the most common reason people believe they have cleaned a document and have not.
What are the five things you should not tell ChatGPT?
That is a data-privacy question, and Google serves it in the People Also Ask box for a text-encoding query. It is out of scope for this page and nothing in this research addresses it, so no answer is offered here rather than an invented one.