Five experienced academics were handed writing by human authors and asked one question: did a machine write this? On the genuinely human text, they got it wrong 72% of the time. Not on the AI samples. On the real ones. They read writing by real people and called it machine-made, nearly three times in four.

The short version: free AI detectors are better at this than people are, and neither is good enough to accuse anyone with. Human raters identified AI text 19% of the time, which is what guessing looks like. Two of three tested detectors tracked AI use closely and agreed with each other at an ICC of 0.96. But those same detectors read non-native English as machine writing at a 61% false-positive rate. If someone has accused you of using AI, the evidence they are holding is weaker than they think it is.

Are AI detectors accurate?

Partly, and the part that works is narrower than the marketing suggests. At the extremes they are good. In a 2026 study of healthcare-simulation articles, ZeroGPT scored genuinely human introductions 6.5 out of 100 and fully AI-written ones 92.5. PhraslyAI scored 5.9 and 92.4. Both moved in the right direction through every stage of AI involvement. The problem is not the extremes. It is whose human writing you hand them.

The study, published in PMC, took 30 open-access articles from two healthcare simulation journals and pulled their introductions — 500 to 600 words each, all written before 2022, so none of them could have been AI-assisted. Then it produced five versions of each: untouched human, lightly AI-edited, heavily AI-edited, AI-written from human bullet points, and fully AI-written from the title alone. Three free detectors scored every version.

ConditionZeroGPTPhraslyAIGrammarly
100% human6.55.91.6
Light AI edit20.224.83.0
Heavy AI edit43.145.111.5
AI from bullet points89.985.662.5
100% AI92.592.450.0

Mean AI score out of 100, by condition and tool.

Read the bottom two rows of the Grammarly column. It scored AI-from-bullets at 62.5 and fully AI text at 50.0 — it rated the more machine-written text as less machine-written. That single inversion is what drags the cross-tool agreement down: ZeroGPT and PhraslyAI agree with each other at ICC 0.96, but each agrees with Grammarly at only 0.57 to 0.60. So "AI detectors disagree with each other" is not quite the finding. The finding is that two of these three behaved consistently and one did not, and from the outside you cannot tell which one you are using.

How well can a person spot AI writing?

They cannot. Across all five conditions the human raters were right 19% of the time, which the authors describe as indistinguishable from chance. Broken down: 17% on untouched human text, 23% on light edits, 30% on heavy edits, 17% on AI from bullets, and 10% on fully AI-written text. They were at their worst on the samples that were entirely machine-written.

The error rates are the part worth sitting with. On text written entirely by humans, the raters produced a 72.2% false positive rate. On text written entirely by AI, a 76.9% false negative rate. And the agreement between the human raters and each of the three tools was ICC = 0 — not weak agreement, none.

These were not random readers. They were healthcare simulation educators and researchers with between one and 25 years of experience, reading articles from their own field, where they would know the conventions better than almost anyone. That is the ceiling on human judgement here, not the floor.

Why detectors flag writers whose first language is not English

Because they are largely measuring how predictable your word choices are, and writing in a second language makes your word choices more predictable. In 2023, Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou ran seven widely used GPT detectors over 91 TOEFL essays by non-native English writers and 88 essays by US eighth-graders. On the eighth-graders, the detectors were close to perfect. On the TOEFL essays, the average false-positive rate was 61.22%, and all seven detectors unanimously flagged 18 of the 91 essays — 19.78% — as machine-written.

The mechanism is perplexity: a measure of how surprising each next word is. Someone writing in a language they learned second reaches for the common, safe word, because that is what working carefully in a second language looks like. Low surprise reads as machine.

The researchers then proved it was style and not substance. They had ChatGPT rewrite the same TOEFL essays with richer vocabulary — same authors, same arguments, same essays — and the average false-positive rate fell from 61.3% to 11.6%. Nothing about the authorship changed. Only the vocabulary did. Their paper, "GPT detectors are biased against non-native English writers", ran in Patterns, and it ends by arguing these tools should not be used in evaluative or educational settings at all.

I have a personal stake in this one. English is not the language I think in first. When a tool scores my writing as machine-made, I have no way to know whether it is detecting a machine or detecting me.

What to do if you have been accused of using AI

Ask four questions, in this order.

  1. Which tool, and what score? A screenshot of a percentage is not a finding. Different tools disagree with each other on the same text, and one of the three tested above rated fully AI writing as less machine-like than partly AI writing.
  2. What is that tool's false positive rate? If whoever is accusing you cannot answer, they are reading a number they do not understand. A tool that is right about AI text and wrong about a large share of human text is not evidence against a specific person.
  3. Does the accusation account for how I write? If you write in a second language, the published research says the tool is biased against you by a measured, published margin. That belongs in the conversation.
  4. Can I show process instead of prose? Drafts, version history, research notes, timestamps, the messy middle. Process is checkable. A style score is not.

One thing not to do: rewrite your work to beat the tool. Raising your perplexity on purpose is exactly the bypass the Stanford team demonstrated, and it works — which is the clearest proof available that the tool is scoring vocabulary rather than authorship. You would be making your writing worse to satisfy an instrument that has already been shown not to measure the thing it claims to.

How we use AI detectors at Skryvo

As a delta, never as a verdict. A draft that scores 95% before an editing pass and 30% after tells you the pass landed. The same score read as a judgement on a finished page tells you nothing you can act on, and chasing a lower score is its own trap. What makes AI writing recognisable is a set of countable habits — uniform sentence rhythm, watermark vocabulary, symmetrical structure — and those are worth fixing because they make writing worse, not because a detector objects to them.

The check we actually gate on is whether the claims are true. When we ran 20 AI-written posts through our fact-check pass, six in ten checkable claims could not be traced to any source. That question has a right answer you can go and verify. "Did a machine write this" does not, and two studies now say nobody — person or tool — can answer it reliably enough to act on. Google's own position points the same way: its guidance rewards people-first content however it was produced, which is why how to build the author E-E-A-T signals that rank AI-assisted content.

What this evidence does not prove

Both studies are small, and we would want this stated by anyone quoting them at us.

  • The detection study covered 30 articles from a single field, scored by five raters using three free tools, with text generated by one model over four weeks in 2025. It is a signal, not a law.
  • The bias study dates from 2023 and tested the detector generation available then. Detectors have changed since; whether the bias has is not something either paper can tell you.
  • Neither tested the paid institutional detectors used in universities and newsrooms, which are the ones most accusations actually come from.
  • TOEFL essays are a specific genre written under exam conditions. They are not a stand-in for all non-native English writing.

What survives all of that: no published evidence supports using any of these scores to accuse a named person, and a good deal of published evidence says the people most likely to be wrongly accused are the ones writing in their second language.

The check that actually works

Stop asking whether a machine wrote it. Ask whether it is true. Run your own draft through a fact-check pass and see how many of its claims survive contact with a primary source — that number is real, it is yours, and unlike a detector score, you can do something about it. It is free to try and there is no card required. The broader question is whether a page is slop at all — and the four-signal test for AI slop checks what is absent, not how it was typed.

About the author
Photo of Muhammad Shakil

Muhammad ShakilFounder of Skryvo

I wrote this myself. I have been building apps for clients since 2019 — Flutter developer, agency founder, Top Rated Plus on Upwork — and I built Skryvo after one too many AI drafts fell apart the moment someone checked a fact. Everything I publish here is tested on my own work first, and I show the numbers, including the ones that did not go my way.

Write content you can publish with confidence

Skryvo generates, fact-checks, and scores your content — try it free.

Start for Free