AI Detection
What an AI detector score actually means
What a detector percentage measures, where our tool draws its verdict lines, and why short texts and non-native writing get misread.
A detector score is an estimate about a piece of text. It can’t tell you who typed the words, whether someone used AI for an outline and wrote the rest themselves, or whether one sentence was pasted in from a chatbot. It reports how closely the writing matches patterns a model learned from human and machine text.
What the number is based on
Detectors are trained on examples of both kinds of writing and pick up what tends to separate them: how predictable the word choices are, how much sentence length varies, how often the same structures come back. A new text is compared against those patterns. The result is a percentage, and the same percentage can mean different things in different tools, because each tool decides for itself where “likely AI” begins.
Here is where ours draws the lines:
- It scores each passage and the text as a whole. The final estimate is 70% the passages (weighted by length) and 30% the whole text.
- It says “Likely AI-generated” only when the AI estimate is 90% or higher.
- It says “Likely human-written” when the AI estimate is 30% or lower and the human estimate is at least 70%.
- Everything else is “Inconclusive”. So is any result the model isn’t confident about, including texts where the passages and the whole text disagree sharply.
A lot of real writing lands in that middle band. Some tools don’t show a middle category at all and report every result as a lean one way or the other. Before you act on a score from any detector, find out where its thresholds are.
Read the flagged sentences
The percentage summarizes. The highlighted sentences are where you can check it. Read what was flagged: sentences of nearly the same length, “Moreover” and “Furthermore” opening paragraph after paragraph, each paragraph ending by restating its point. Then ask whether those habits could be the writer’s own. Formal assignments, templates, and technical or legal writing often look like this with no AI involved.
Short texts are hard to judge
Our detector won’t scan anything under 50 characters, and even a paragraph gives it little to go on. A title, a few bullet points, or a 100-word answer can tip one way on a single phrase. Scan the whole document instead of an excerpt when you can.
Who gets flagged wrongly
False positives don’t land evenly. In a 2023 Stanford study, several widely used detectors consistently labeled essays by non-native English writers as AI-generated, while essays by native speakers were classified correctly (Liang et al., Patterns). On average, more than half of the non-native essays were flagged. The likely reason is that a smaller vocabulary and more predictable phrasing look “machine-like” to these models.
Using the result
Treat a score as a reason to look closer. If authorship matters, ask for drafts or version history, and talk to the writer about how the piece came together.