About the AI Text Detector
Every other tool on this search returns a percentage. This one returns the measurements that percentages are built from, and stops there. That is not modesty. Publishing a score means publishing an accuracy figure, and an accuracy figure nobody has independently measured is an invented number. What follows is what the detectors measure, what they publish about their own error rates, and how far apart those published numbers are.
What does this tool measure?
Five signals, each reported as a number with the sample size it came from. Sentence length variation, which is the standard deviation of sentence length divided by the mean. Em dash density per thousand words. The count of typographic quotes and apostrophes. The count of invisible Unicode characters. And matches against a list of thirty-five phrases common in machine drafting.
Each figure appears beside a sentence describing what it cannot establish, because that sentence is the part most readers need and the part most tools omit. Nothing is uploaded and nothing is stored; the arithmetic happens in the page you are reading.
Why does this tool not give you an AI percentage?
Because a percentage is a classifier output, and a classifier has to publish an accuracy figure. We have not independently measured one for this page, so printing a number next to the word accuracy would be inventing it. Nine of the tools ranking for this search return a score from zero to one hundred. This one shows the measurements those scores are computed from and lets you read them.
The trade is deliberate. A score is easier to act on and harder to check. Five numbers are harder to act on and trivial to check, and on a question where acting wrongly means accusing a person of dishonesty, being checkable is worth more than being convenient.
How do AI detectors decide?
They score how predictable the text is. QuillBot describes its own approach as evaluating how predictable the text is, the variation in sentence structure, and repetitiveness. Those three things are measurable, and two of them are on this page.
Sebastian Raschka, a machine learning researcher whose from-scratch detector build ranks on this same search, lists the wider family: supervised classifiers, perturbation-based probability tests, perplexity measures and watermarking. His post is partly behind a paywall, and the free portion carries the structural point the vendor pages leave out: a new model may not exhibit the pattern a detector was trained on, so the detector has to be updated to catch it, and then updated again. Detection is a chase rather than a solved measurement.
How accurate are AI detectors, according to the detectors?
Their own published figures, side by side, are the fastest way to see the problem. Copyleaks publishes over 99% accuracy. Pangram publishes 99.98% accurate with third-party verified results. QuillBot publishes a 99% detection rate, and credits it to independent evaluation on RAID, a named third-party benchmark, which is the only claim in this set that points somewhere a reader can go and check.
ZeroGPT's page carries no accuracy figure at all in its rendered text. Its FAQ contains the question about its accuracy rate as a collapsed panel whose answer did not load when we captured it, so we record that as not captured rather than as absent.
Now the part the figures hide: detection rate, accuracy and false positive rate are three different measurements, and these pages use them as if they were interchangeable. A high detection rate means machine text is usually caught. It says nothing about how often human text is wrongly flagged, which is the number that decides whether a person gets accused.
What is a false positive rate, and whose is published?
A false positive is human writing flagged as machine written, and four of the ranking detectors publish a rate for it. They disagree by a factor of one hundred.
Copyleaks publishes an industry-low 0.03% false positive rate, in its own product copy. Pangram's figure of 1 in 10,000 is more specific still, and it appears inside a testimonial from a named professor rather than in Pangram's own voice, which is a distinction worth keeping. GPTZero publishes 1% for ESL writing. Paperpal says its multi-model checks reduce false positives by over 40%, which is a relative improvement with no baseline stated anywhere on the page, so it cannot be checked: a drop from 10% to 6% and a drop from 0.05% to 0.03% are both over 40%.
And QuillBot, the first result on this search, carries an FAQ headed with the question of its own false positive rate. The answer contains no number. It is a careful, sensible answer about the risk of false positives in general, and it does not state a rate.
Are non-native English writers flagged more often?
GPTZero is the only page in this set that publishes a false positive rate for a named group, and the group it names is the one most at risk. In its own words: it is the only AI detector de-biased for ESL learners, which many competitors do not take into account, and it repeatedly trains its model to reduce the false positive rate for ESL writing to 1%.
Credit where it belongs: publishing that figure is better behaviour than not publishing it, and the number looks bad precisely because it is honest about a hard case. One percent is thirty-three times Copyleaks' headline 0.03%, which is the comparison the category never makes.
What cannot be said, and we will not say it: that the other detectors are worse for this group. They do not publish a figure for it. An absent number is an absent number, not a bad one.
What does a 1% false positive rate mean at scale?
It means about one hundred people, per ten thousand submissions. That is arithmetic on GPTZero's own published figure for ESL writing, not an estimate of ours. A department screening ten thousand pieces of coursework from second-language writers, at a 1% false positive rate, produces roughly a hundred students whose own writing is flagged as machine written.
Run the same ten thousand through Copyleaks' published 0.03% and you get about three. Both numbers describe the same product category, and which one a reader believes depends entirely on which page they landed on. Neither figure is ours, and we have not tested either.
Does careful, formulaic writing get flagged?
Yes, and the market leader says so in its own documentation. QuillBot states that formulaic human writing, naming academic definitions, legal copy and structured templates, can score higher than expected for any AI writing detector.
The consequence follows and is worth stating: the more disciplined and conventional the human writing, the closer it sits to the distribution these tools flag. A student writing carefully to a rubric, in a genre with fixed conventions, is more exposed than one writing loosely. That is the opposite of how detection results are usually read.
How much text does a detector need?
Far more than the paragraph most people paste in. QuillBot requires a minimum of 80 words to scan at all, and states that texts of 300 words and above produce more reliable scores than short inputs. GPTZero states that accuracy improves with longer inputs, that document-level results are stronger than paragraph or sentence level, and that it is strongest on English prose.
So the short excerpt that triggered someone's suspicion is exactly the input on which every tool in this category is least reliable. Two vendors publish that, and it is not the sentence anyone quotes. This page warns you when a sample is under a hundred words, because below that the variation figure is noise rather than a measurement.
What is burstiness, and what does it measure?
It is the variation in sentence length across a passage, and on this page it is reported as a coefficient rather than a mood. Take the standard deviation of the sentence lengths, divide by the mean, and you have the figure. Human prose tends to swing between short sentences and long ones; machine prose tends to cluster nearer its own average.
We report the mean, the standard deviation and the sentence count alongside it, so the figure can be recomputed by hand. Two failure modes deserve equal billing: text that a person edited heavily and text a person wrote evenly both land in the middle of the range, and a short sample produces a number that means nothing at all.
Is an em dash evidence that a machine wrote it?
No. It is the most repeated surface tell and one of the weakest. We report density per thousand words rather than a raw count, because a raw count mostly measures how long the passage is.
Frequent em dashes are ordinary in edited human prose, and word processors manufacture them automatically from two hyphens, so the character enters documents that no model ever touched. A high density is worth noticing and settles nothing on its own. If you have measured it and want the character gone, the em dash remover handles the replacement and the spacing it leaves behind.
Do invisible characters prove AI?
No, and we measured this rather than assuming it. Zero-width characters and other hidden codepoints are artefacts of the pipeline the text travelled through, not signatures a model applies.
We scanned 90,880 words of Claude output for every class of hidden character this tool tracks and found none at all. That measurement, and the mechanism behind the watermark people expect to find there, are on the Claude watermark remover, which also removes the characters if a document does turn out to carry them. If you want to inspect them rather than strip them, the invisible character detector explains what each one is for.
Is a phrase list evidence?
It is the crudest signal on this page and it is labelled that way in the output. We match against thirty-five phrases that turn up often in machine drafting, and we show you which ones matched rather than only how many.
Every phrase on that list is also written by people, constantly, and several of them are simply conventional English connective language. A high count says the passage is written conventionally, which correlates with genre and with training rather than with authorship. The list is a heuristic we chose. It is not a standard, and nobody should treat it as one.
How is this different from a plagiarism checker?
They answer unrelated questions. A plagiarism checker asks whether this text appears somewhere else already. A detector asks whether the patterns look machine generated.
Writing can be entirely original and entirely machine written, and it can be copied word for word from a human author with no model involved anywhere. Neither tool answers the other's question, and a clean result from one says nothing about the other.
Can you use this to accuse someone?
No, and the detectors themselves say the same thing. Nothing measurable in a piece of text establishes who wrote it. Every signal here is consistent with several explanations, and the passage in front of you cannot tell you which one applies.
QuillBot's own documentation states that results are probability estimates rather than verdicts, that its detector does not verify human originality, and that nobody should rely on AI detection alone for a decision affecting someone's career or academic standing. The clearest sentence in the whole category belongs to a customer rather than a vendor: a professor quoted on Pangram's site describes a flag as an invitation to talk, never a presumption of guilt. That is the correct use of every number on this page.
What this tool will not do
It returns no score, no percentage and no verdict, and it has no accuracy figure because it does not classify anything.
It does not process images, video or files. It has no API. It stores nothing and uploads nothing. Its measurements assume English prose: sentence length variation and an English phrase list do not carry over to other languages without testing we have not done, so the tool does not claim multilingual support. And it carries no advice on avoiding detection, in either direction.