How accurate is our AI detector?
Last tested 28 September 2026 on 336 human-written and 698 AI-written texts.
An AI detector gives a reading, not proof, and a reading is only as good as its error rates. Here are ours, measured on writing whose author we know, and the limits you should hold us to. We update this page when we retest.
What each result means
How the texts in our test landed in each of the three results the detector can give.
| Result | Human-written336 texts | Written by current AI358 texts |
|---|---|---|
| No AI signal foundScore 0 to 25 | 85%285 of 336 | 39%139 of 358 |
| BorderlineScore 26 to 60 | 15%51 of 336 | 34%123 of 358 |
| Reads like AIScore 61 to 100 | 0%0 of 336 | 27%96 of 358 |
None of the 336 human-written texts read “Reads like AI”. That is the number we protect first, because a false flag can put a student in front of a misconduct panel. About 15 in 100 landed in Borderline, which is why we call that band inconclusive rather than a warning.
On older, more formulaic AI writing it does better: 45 in 100 of those texts read “Reads like AI”. Today’s models write more cleanly, which is why the current-model column is the one to trust.
What it cannot do
It cannot prove who wrote something.
No detector can. A reading describes patterns in the text, not where the text came from. Treat any score, ours included, as one signal and never as a finding about a person.
It misses a lot of writing from current AI models.
In our test it found no AI signal at all in 39 in 100 texts written by current AI models. A low score is not clearance, and we say so under every low result.
It rarely catches AI text that has been rewritten or heavily edited.
Rewriting and editing remove most of the patterns a detector looks for, so edited AI text usually reads as human to it.
Research writing is its weakest ground.
Dense academic prose, like a research abstract, is the human writing detectors misread most. 31 of the 79 research passages in our test read Borderline, so for that kind of writing we need stronger evidence before we say "Reads like AI".
The same text can move a few points between checks.
We checked 23 texts three times each. The typical score moved 1 point, and the most any text moved was 8. A score near the edge of a band belongs to both.
It is built for English.
Text written mostly in another alphabet is not scored. We have not tested other languages written in the Latin alphabet, so we make no claim about them.
When our full check cannot run, you get no number.
If the full check is unavailable, the result says “Not scored” and asks you to try again, instead of a guess from a weaker method.
How we tested
Every text went through the same check you use on the detector page, on 28 September 2026. We only used writing whose author we know:
- 336 human-written texts: passages from research papers posted to arXiv between 2015 and 2019, Wikipedia articles as they stood before 2022, public-domain books from Project Gutenberg, and answers on public question-and-answer forums.
- 358 texts written in 2026 by three different AI providers. Each one was written on the same topic, in the same kind of writing, as one of the human texts, so the comparison is like for like.
- 340 texts from older AI models, across 34 kinds of writing from essays to cover letters.
Some kinds of writing appear in the test only 10 to 60 times. Until we have tested at least 300 texts of one kind, we do not claim a rate for that kind on its own.
Where “Reads like AI” starts
We set that line by a rule, not a round number: no more than 1 in 100 honest texts may land above it. When we tested our old line (55), 6 of 336 human-written texts did, all of them research abstracts. So we moved the line to 60 and now ask for stronger evidence before dense research writing can cross it. At the new line none of the 336 did, including when we checked the closest ones twice more.
The cost is real: the top band now catches less AI writing than before. We think that is the right trade for a result students may have to answer for.