Research
Are AI detectors accurate? What the research shows
Vendor accuracy claims, what independent research found, why false positives matter more than they look, and what our own testing showed.
Most AI detectors advertise accuracy of 99% or close to it. Those figures generally come from the vendor’s own tests, on texts and at cutoffs the vendor chose. Independent studies have found more errors than that, and the errors don’t fall on everyone equally.
What a single accuracy number leaves out
A detector can be wrong in two ways. It can call human writing AI (a false positive), or it can miss AI writing (a false negative). One accuracy figure blends the two, and for most uses they aren’t equally bad. Missing some AI text is a nuisance. Telling a student who wrote their own essay that a machine wrote it is worse.
False positives also add up with volume. Suppose a detector wrongly flags 1% of human writing, which would be a good result. A school that scans 5,000 essays written without AI would still get about 50 flags, each one a student who did nothing wrong. The 1% is an assumption to show the arithmetic, not a measured rate for any tool.
What independent research found
The most cited study is from Stanford: Liang and colleagues, published in Patterns in 2023 (paper). They ran seven widely used detectors on 91 essays that non-native English speakers wrote for the TOEFL exam. On average, the detectors labeled 61.3% of these human-written essays as AI-generated. At least one detector flagged 97.8% of them, and 19.8% were flagged by all seven. Essays by native English speakers were classified accurately.
The researchers point to how detectors work: they lean on how predictable the wording is, and writers with a smaller English vocabulary tend to choose more predictable words. The same paper showed that simple prompting could remove the bias, and could also get AI-written text past the detectors.
Detectors have changed since 2023, and some vendors now publish much lower false-positive rates. Those are worth checking against independent tests where they exist.
What we found testing our own detector
We ran a small pilot on the HC3 dataset, which pairs human and ChatGPT answers to the same questions. On a held-out sample of 20 human answers and 20 ChatGPT answers, the number of human answers flagged depended mostly on where we drew the line:
| Score counted as “AI” | Human answers flagged | ChatGPT answers caught |
|---|---|---|
| 70% or higher | 5 of 20 | 20 of 20 |
| 90% or higher | 0 of 20 | 20 of 20 |
An earlier version of our product called anything at 70% or higher “Likely AI-generated”. On this sample, that would have wrongly flagged a quarter of the people. It now waits for 90%, and reports the middle range as “Inconclusive”.
Twenty examples of older ChatGPT output is a small sample, so read this as a point about cutoffs, not an accuracy claim. Zero false positives in 20 does not mean zero in real use.
Using a detector without overtrusting it
- Scan complete documents. Short passages give any model little to go on.
- Find out where the tool’s cutoffs are, and whether it has an inconclusive category.
- Read the highlighted sentences instead of stopping at the percentage.
- Take extra care with writing by non-native English speakers and with formal or formulaic genres.
- Pair any score with evidence about how the piece was written, such as drafts or version history.
If you’re on the other side of a flag, what to do if you’re falsely flagged covers the steps from the writer’s side.