Paste the same 600-word article into three AI detectors and you can easily get three different answers: 12% AI from the first, 71% from the second and “uncertain” from the third. It is tempting to conclude that one of them is broken, or that all of them are guessing.
Neither is quite true. AI detectors disagree because they were built from different data, they report different kinds of numbers, and they make different decisions about where to draw the line. Once you know where those differences come from, a disagreement stops being confusing and starts telling you something useful.
They learned from different writing
Every detector is trained on examples of human writing and AI writing, and it learns what “normal” looks like from whatever it was shown. If a detector’s human examples were mostly news articles and student essays, a casual product review or a technical how-to may look unfamiliar to it. Unfamiliar can look suspicious.
The AI side matters just as much. Chatbots change quickly, and a detector trained mostly on text from older models may be sharper on those models than on the ones people use today. Two detectors with different training sets can look at the same paragraph and see different things.
They don’t report the same kind of number
A result of “40%” sounds universal, but it is not. Depending on the tool, a percentage might mean any of these:
- the share of the document the tool thinks was machine-written;
- a confidence-style estimate for the text as a whole;
- a score on the tool’s own scale that is neither a share nor a probability.
So 40% in one detector and 40% in another are not necessarily the same finding, and 40% in one against 70% in another is not necessarily a contradiction. Before comparing results, read what each tool says its number means.
They draw the line in different places
Every detector has to decide how much evidence is enough to say “AI”. Set the bar low and it catches more machine text but wrongly flags more people. Set it high and it flags fewer people but misses more AI. There is no setting that avoids both, so each company picks its own trade-off.
Some tools are open about this. Turnitin’s guidance, for example, explains that it shows an asterisk instead of a score for AI results between 1% and 19%, because its testing found a higher incidence of false positives in that range. Another tool might show those same low scores as numbers, which makes it look more “suspicious” without actually knowing more.
They treat short and formatted text differently
Length is one of the biggest reasons results jump around. A short passage gives a detector very little evidence, and small changes in wording can swing the score. OpenAI said its own classifier, which it withdrew in 2023, was “very unreliable” on texts below 1,000 characters. Turnitin asks for at least 300 words of prose in a long-form format before it produces an AI report at all.
Other tools will happily score a 40-word paragraph, and that result is close to noise. Detectors also differ in what they do with headings, bullet lists, quotations, references and code. Some strip these out, some score them, and a few quietly skip anything that isn’t continuous prose.
They behave differently outside plain, native English
Most detectors were built and tested mainly on English, and writing from non-native speakers can trip them up. In a 2023 Stanford study, seven detectors classified more than half of 91 TOEFL essays written by non-native English speakers as AI-generated (61.22% on average), and 89 of the 91 essays were flagged by at least one detector. If the text is in another language, or written by someone using English as a second language, expect detectors to disagree more.
Even one detector can change its answer
It isn’t only different tools that disagree. The same detector can give a different result for what looks like the same text:
- Paste versus upload. Uploading a PDF or Word file can pull in headers, footers, captions and references that were not in the pasted version.
- Small edits. On a shorter text, rewriting a few sentences can move the score a long way, because each sentence carries more weight.
- Model updates. Detectors are retrained and re-tuned. A text checked in spring and again today may be scored by a different version of the same tool.
If a result matters, note the date, the exact text you checked and how you submitted it, so the check can be repeated.
How to decide which result to trust
You can’t make every detector agree, but you can weigh their answers sensibly.
Prefer tools that publish their error rates
The most useful number a detector can give you is how often it wrongly flags human writing, measured on text it never saw during training. Few tools publish it. One free AI detector, for instance, shows the false-flag rate for your text’s length next to every score: in its sealed test, about 2 in 1,000 human texts of 150 words or more were wrongly flagged, and about 2.4 in 100 at 100 to 149 words. A tool that tells you how often it is wrong gives you something to reason with. A tool that only gives you a big red percentage does not.
Prefer ranges and bands over exact percentages
A score of 73% looks precise, but the real uncertainty behind any single number is wide. Tools that show a range, or sort results into bands such as “likely human”, “uncertain” and “likely AI”, are being more honest about what they can and can’t tell.
Check the longest continuous writing you have
If you only have a short answer or a single paragraph, don’t check it. Find the longest stretch of continuous prose by the same writer, ideally a few hundred words, and use that.
Use more than one detector, and read the pattern
Detectors built on different data make different mistakes. When two or three of them agree on a long text, that agreement is worth more than any single score. When they split, the honest reading is that the evidence is mixed.
Never treat a score as proof
Even the companies that build these tools say so. OpenAI withdrew its own classifier in July 2023 because of its “low rate of accuracy”; at launch, it had correctly labelled only 26% of AI-written text as likely AI and wrongly flagged human text 9% of the time. Turnitin states that its AI score should not be used as the sole basis for adverse actions against a student. A score is a reason to look closer, not a conclusion.
A quick guide to reading mixed results
| What the detectors say | What it usually means |
| All low | Probably fine. There is nothing here that needs acting on. |
| All high on a long text | Worth a closer look. Check drafts, sources and version history before drawing any conclusion. |
| Split results | Treat it as uncertain. Disagreement between differently built tools is itself the answer: the evidence is mixed. |
| Any result on a very short text | Don’t decide anything. Find a longer sample of continuous writing and check that instead. |
The bottom line
AI detectors disagree because they are measuring slightly different things, trained on different writing, and tuned to different levels of caution. That makes any single score less meaningful than it looks, and a pattern of results far more meaningful. Choose tools that tell you how often they are wrong, check long passages instead of snippets, compare more than one result, and keep the final judgement with a person who has read the text.