AI Detection

How AI detectors work, and why they get things wrong

The three ways AI text detection works (predictability, trained classifiers, and watermarks), where each one fails, and how our detector reaches a verdict.

By aitextdetectly team4 min read

An AI detector doesn’t find a hidden label in the text. Most of them look at how the writing is put together and estimate how closely it matches what language models tend to produce. That estimate can be useful, but it’s built from patterns that people produce too, and that’s where most of the mistakes come from.

There are three main approaches. Many tools combine more than one.

1. Measuring how predictable the words are

A language model can score how likely each next word is, given the words before it. Text where almost every word is the likely one has low “perplexity”. Text full of surprising word choices has high perplexity.

Model output tends to be predictable, because generating text means picking likely words. Human writing tends to be less even: a plain sentence, then an odd phrase, then a long winding one. Some detectors measure that variation across sentences too, often called “burstiness”. GPTZero described its early approach in these two terms.

The weakness is plain to see. Plenty of people write predictably. Someone writing in a second language, a student following an essay template, or a lawyer drafting a contract all tend to choose the expected word. To a predictability measure, that looks like a machine.

2. Training a classifier on examples

The second approach trains a model on large sets of text labeled “human” and “AI”, and lets it learn whatever separates them: word choice, rhythm, structure, how paragraphs open and close. A new text gets a score based on which set it resembles more.

This can pick up patterns that predictability misses, but a classifier is only as good as its examples. It can struggle with models newer than its training data, with genres it rarely saw, and with text a person has edited. OpenAI released a classifier of this kind in January 2023. By OpenAI’s own figures it caught 26% of AI-written text and wrongly labeled human text as AI 9% of the time (OpenAI). It withdrew the tool that July, citing its low accuracy.

3. Watermarks added at generation time

A watermark works differently. The company running the model nudges its word choices in a pattern that’s invisible to readers but can be checked statistically later. Google DeepMind published a method of this kind, SynthID Text, in Nature in 2024 (Dathathri et al.).

A watermark is the strongest evidence of the three when it’s there, but it’s narrow. It only exists in text from a model that applies one, only the right checker can read it, and paraphrasing or translating the text can weaken it. It says nothing about text from any other model.

Why detectors get things wrong

The errors cluster in predictable places:

  • Predictable human writing. In a 2023 Stanford study, seven widely used detectors labeled 61% of essays by non-native English speakers as AI-generated, on average, while essays by native speakers were classified correctly (Liang et al., Patterns).
  • Short texts. A few sentences give any method too little to measure.
  • Edited text. A human editing an AI draft, or an AI tool polishing a human draft, produces something in between, and the score lands in between too.
  • Newer models. A detector tuned on older output can miss the habits of newer models, or mistake new human styles for them.

A single percentage hides all of this, which is why it matters where a tool draws its lines.

How our detector reaches a verdict

We use a language-model classifier from our detection provider, TypeSafe, and ask it to judge the text three ways: the whole text, passages of about 50 words, and individual sentences. The AI, mixed, and human estimates are 70% the passages, weighted by length, and 30% the whole text. The sentence judgments drive the highlights, so you can see which parts the score comes from.

The verdict says “Likely AI-generated” only at 90% or higher, and “Likely human-written” only when the AI estimate is 30% or lower and the human estimate at least 70%. Everything else, and any result where the model isn’t confident or the passages and the whole text disagree sharply, is “Inconclusive”. The About page has what our own testing found and its limits.

Reading any detector’s result

Treat a score as a reason to look closer, not a finding. Read the sentences it flagged, find out where its cutoffs are, and if authorship matters, look at how the document was written: drafts, notes, and version history. What an AI detector score actually means covers reading a result in more detail.