Skip to content
appvior
← All guides

Guide

How accurate are AI detectors? We tested twenty texts of known origin

Every detector publishes an accuracy figure and none publishes a test you can repeat. So we ran one. Ten texts published before ChatGPT existed, ten written by an AI model on the day of the test, all matched by genre, through two free detectors. The two tools disagreed with each other on exactly half the samples, and the most confidently flagged human text was written in 1845.

The headline result

Neither free detector was reliable, and they failed in opposite directions. Against twenty samples where a coin would score 50%, ZeroGPT got 11 right and Sapling got 13.

  • ZeroGPT — flagged 3 of 10 genuine human texts as AI, and missed 6 of 10 AI texts. 11 of 20 correct.
  • Sapling — flagged 6 of 10 genuine human texts as AI, and missed 1 of 10 AI texts. 13 of 20 correct.
  • The two tools returned opposite verdicts on 10 of the 20 samples — every second text.
  • A sample counted as "called AI" at 50% or above. Both tools report one percentage and neither publishes an action threshold, so the line is ours and we are stating it.

What we tested, and why it is checkable

Ten human samples all published before ChatGPT existed, so their authorship is not a matter of opinion, and ten AI samples written the day of the test in the same genres.

The human set: the opening of Pride and Prejudice (1813), a chapter of Frederick Douglass's narrative (1845), Emily Dickinson's poems (published 1890), five academic abstracts submitted to arXiv and PubMed in 2018 and 2019, and two Wikipedia extracts pulled from named May 2019 revisions by permanent revision ID. The AI set: a novel opening, a memoir passage, a poem, four academic abstracts in the same fields, a structured medical abstract, and two encyclopedia articles on the same topics as the Wikipedia extracts — all written by Claude Opus 5 on 17 August 2026, one attempt each, unedited. Samples ran 155 to 309 words. Genre matching matters: a test that pits Victorian prose against modern chatbot output measures the century, not the detector.

The false positives, named

Six human texts were flagged, and the pattern in them is the opposite of reassuring: it is edited, conventional, well-structured prose that trips these tools.

  • Frederick Douglass, Narrative of the Life of Frederick Douglass (1845) — 99.9% AI on Sapling.
  • Emily Dickinson's poems (published 1890) — 99.8% AI on Sapling. This is the plain answer to "the AI detector says my poem is AI": compressed, regular, patterned language is what these models score as machine-like.
  • Wikipedia's photosynthesis article as it stood in May 2019 — 100% on ZeroGPT and 99.1% on Sapling. The only human sample both tools flagged, and it predates generative models being available to write it.
  • Wikipedia's Great Depression article, May 2019 — 72.7% and 99.3%.
  • The opening of Pride and Prejudice — 61% AI on ZeroGPT, while Sapling gave the same passage 12.5%.
  • A 2019 computational linguistics abstract and a 2018 neuroscience abstract, both real, both flagged by Sapling at 99.9% and 59.3%.

The misses, which matter just as much

ZeroGPT scored four AI-written academic abstracts at 0.0% — the lowest possible score — while Sapling scored three of those same four at 100%.

The most instructive single result: an AI-written structured medical abstract scored 0.0% on Sapling and 23.6% on ZeroGPT, while the genuine 2018 medical abstract it was modelled on scored 0.1% and 0.0%. Both tools treated the fabricated one as no more machine-like than the real one. Rigid genre conventions — BACKGROUND, METHODS, RESULTS, CONCLUSION — defeat detection in both directions, because the format constrains the writing enough that human and model output converge.

Why one number cannot settle a case

Because the second tool you run will often disagree, and there is no principled way to decide which one was right.

Half our samples got opposite verdicts. Someone accused on a ZeroGPT score can produce a Sapling score that says the opposite, and vice versa, and neither party can appeal to a published, reproducible validation set to break the tie. This is exactly why Turnitin — the strictest tool in this market and the one with the most detailed public methodology — states that its own percentage should not be used as the sole basis for action. If the vendor with the best claim says the number is not a finding, a free tool's number certainly is not one.

If you have been flagged and you wrote it

Produce the process, not a counter-score. Version history is the strongest evidence available to you, and it exists only if you did not paste your draft in as a single block.

  • Google Docs or Word version history showing the document being built over time.
  • Notes, outlines and earlier drafts, in whatever state they are actually in.
  • Library or browser records of the sources you read.
  • A second detector's score on the same text, which in our test disagreed with the first half the time — useful as evidence of unreliability, not as proof of innocence.
  • The vendor's own limits in writing: Grammarly states that no AI detector is 100% accurate, and Turnitin states that its indicator is not a determination of misconduct.

What about Grammarly's detector, and Turnitin's

Neither is in our results, for different reasons, and we would rather say so than pad the table.

Turnitin cannot be tested without an institutional licence, so this measures free consumer tools rather than the detector most colleges actually run. Grammarly's free detector we tried to include and could not measure reliably: our first readings turned out to be stale — the results panel kept displaying the previous score while the editor silently rejected new input — and by the time we caught that, the site was asking us to create an account rather than run further scans, which we do not do on services we are testing. Grammarly's published claim is 99% detection accuracy and a first-place ranking on the RAID benchmark; that is its claim, not our finding, and its own page also states that no AI detector is 100% accurate.

The limits of this test

Twenty samples, two detectors, one generating model, one day. That is enough to show these tools contradict each other constantly. It is not enough to put a decimal point on anyone's false positive rate.

  • Two free detectors, not the institutional tools. Turnitin, Copyleaks and GPTZero's paid tiers were not tested.
  • All ten AI samples came from one model. A detector tuned on other models' output may score them differently, in either direction.
  • One run per sample per detector, so run-to-run variance is not measured.
  • The 50% threshold is ours. Both tools' scores were heavily bimodal, so most samples sat far from the line, but a different threshold changes the counts.
  • These are the models running on 17 August 2026. Detectors update frequently, and a repeat of this test in six months should be expected to produce different numbers.

Disclosure

We make an app that scores text for AI-likeness and rewrites it, so we have an interest here, and it cuts both ways: this page argues that scores of this kind — including the one our own app reports — are weak evidence.

The corpus and every per-sample score are published on this site, ungated, so the test can be repeated or contradicted. If you rerun it and get different numbers, that is the point of stating the method. What we will not do is publish an accuracy figure for our own detection score alongside these, because a vendor grading itself on its own corpus is the practice this page exists to criticise.

Questions

Common questions

How accurate are AI detectors?

In our 17 August 2026 test of twenty texts of known origin, ZeroGPT was right 11 times out of 20 and Sapling 13 out of 20, where chance is 10. They returned opposite verdicts on half the samples. Institutional tools were not tested; Turnitin claims under 1% false positives on documents over 20% AI writing.

Do AI detectors flag human writing?

Yes, and often. Six of our ten pre-ChatGPT human samples were flagged by at least one detector, including Frederick Douglass's 1845 narrative at 99.9%, Emily Dickinson's poems at 99.8%, and Wikipedia's 2019 photosynthesis article at 99.1% and 100%.

Why does an AI detector say my poem is AI?

Because poetry has the properties these models score as machine-written: short, regular lines, compressed phrasing and consistent structure. Emily Dickinson's published poems scored 99.8% AI in our test. A high score on a poem is close to meaningless.

Is ZeroGPT a good AI detector?

It was the more conservative of the two we tested — 3 false positives out of 10 human texts — but it missed 6 of 10 AI texts, including four AI-written academic abstracts it scored at 0.0%. Low false positives bought at the price of catching little.

Turnitin says I used AI but I didn't. What do I do?

Gather process evidence rather than arguing about the score: version history, drafts, notes and sources. Turnitin's own position is that the indicator should not be the sole basis for action, which is worth quoting because it is the vendor's position rather than yours. Where a formal appeal route exists, use it.

Can two AI detectors give different results on the same text?

Routinely. In our test the two tools disagreed on the verdict for exactly half the samples — one scored an AI-written economics abstract at 0.0% and the other at 100%.

Can I see the data behind this test?

Yes, both files are on this site and neither is gated. Every sample with both scores and both verdicts is at /data/ai-detector-test-2026-08-17.csv, and the full corpus, including the text of all ten AI samples, is at /data/ai-detector-test-corpus-2026-08-17.json. The ten human samples are identified by source and, for the Wikipedia extracts, by permanent revision ID, so they can be retrieved exactly. The AI samples were generated by Claude Opus 5 on 17 August 2026.