Six in ten. That is how many of the checkable claims in 20 AI-written blog posts I could not confirm, even with a fact-checking pipeline built for exactly that job. And the claims that failed were not the vague ones. They were the precise ones: "47% of email recipients open based on the subject line (Mailchimp)", "91.5% of web pages get zero traffic from Google (Ahrefs)".
The short version: AI makes up statistics, and it does it in a recognisable way: it attaches a plausible number to a real, trusted source. In our test, 105 of 186 checkable claims (56.5%) could not be traced to any source, 8 more were contradicted by one, and every one of the 20 articles contained at least one of them. The fix is not avoiding AI. It is never publishing a number you have not traced to its original page.
Does AI make up statistics?
Yes. A language model writes the most likely next words, and in a post about email marketing, "according to Campaign Monitor" followed by a percentage is very likely text. Whether that study exists is a separate question the model is not answering. Ask it for "statistics and expert insights" and it supplies the shape of evidence, with or without the evidence.
This is not a fringe problem. When the European Broadcasting Union and the BBC tested AI assistants on news questions, 45% of answers had at least one significant issue and 31% had serious sourcing problems. In 2025, Deloitte's Australian firm agreed to a partial refund on a government report after reviewers found references to research that does not exist and a misquoted court judgment.
What percentage of AI-generated claims can you trust?
We wanted a number for the kind of content people actually publish, so we ran the test ourselves. The prompt was the one most people use: "Write a 700-word blog post about [topic]. Include statistics and expert insights." Ten everyday business topics (email marketing, conversion rates, SEO for new websites, SaaS retention and six more), run twice, six days apart, on September 18 and 24, 2026, with DeepSeek V4 Flash, the model Skryvo uses as its main writer. Every draft then went through our production fact-check pass, which pulls out each checkable claim and searches for a source that supports it.
| Run 1 (Sep 18) | Run 2 (Sep 24) | Total | |
|---|---|---|---|
| Articles (words) | 10 (9,286) | 10 (9,359) | 20 (18,645) |
| Checkable claims | 88 | 98 | 186 |
| Confirmed by a source | 28 | 45 | 73 (39.2%) |
| No source found | 56 | 49 | 105 (56.5%) |
| Contradicted by a source | 4 | 4 | 8 (4.3%) |
| Articles with at least one of those | 10 of 10 | 10 of 10 | 20 of 20 |
Two details matter more than the headline. First, 110 of the 113 unconfirmed or contradicted claims contained a number. The model rarely invents opinions; it invents figures. Second, 54 of those 113 arrived with an attribution: a named company, "according to", or "a study found". The claims that looked the most sourced were among the least verifiable.
Read the limits before you quote this. "No source found" is not proof a claim is false. It means an automated search could not confirm it, and one of our lookup services timed out during the second run, so some true claims sit in that bucket. It is one model, one prompt and 20 articles: directional, not a census. That is why I traced a sample by hand.
How AI invents a statistic: seven claims we traced by hand
I took seven flagged claims and searched for the original source of each one, the way an editor would.
| What the AI wrote | What the source actually says | Verdict |
|---|---|---|
| "91.5% of web pages get zero organic traffic from Google (Ahrefs)" | Ahrefs published 90.63% in its original study and 96.55% in its current one. Neither is 91.5%. | Real source, invented number |
| "According to a 2023 McKinsey Global Survey, 60% of organizations now use AI in at least one business function" | In McKinsey's 2023 report, 60% is the share of organisations already using AI that had adopted generative AI. One third of respondents used generative AI regularly. | Real number, wrong meaning |
| "Walking meetings boost divergent thinking by up to 81%" | Stanford's 2014 study (Oppezzo and Schwartz): walking raised creativity for 81% of participants on one test; the average gain was 60%. Nobody measured meetings. | Real study, wrong meaning and context |
| "47% of email recipients open an email based solely on the subject line (Mailchimp)" | The figure circulates as a chart on email-vendor blogs. We found no study behind it and nothing from Mailchimp. | Zombie stat, borrowed source |
| "According to Intuit, owners who use AI accounting tools save an average of $4,200 per year" | A search of Intuit and QuickBooks pages turned up nothing matching it. | Untraceable |
| "It takes an average of 23 minutes to refocus after a distraction (UC Irvine)" | Gloria Mark's research found people take about 23 minutes to resume an interrupted task. | Real, slightly reworded |
| "The ROI on email marketing is often $36 for every $1 spent, compared to $2 for paid social" | Litmus does publish $36 for every $1. We found no source for the $2 comparison. | Real headline, invented comparison |
Only two of the seven held up, and both were bent: one reworded, one with an unsourced comparison bolted on. Our automated checker had also missed the Litmus figure, which is exactly why the hand trace matters. Look at the failures: real brands, real studies and real-sounding numbers, recombined. Nothing looks invented, which is why these survive a quick skim.
Why does AI make up facts?
Because it is rewarded for answering, not for knowing. A model predicts text; it has no internal flag that separates "I read this in a study" from "this sounds like something a study would say". Asking for statistics raises the pressure: the prompt says a good answer contains numbers, so it produces numbers. We covered how language models fabricate with confidence — and the workflow for catching it before you publish.
Is AI-generated content accurate enough to publish?
Not by default — but the problem is specific. AI-generated content is accurate enough to publish for structure, explanation and first drafts. It is not accurate enough to publish when it contains a number, a date, a quote or a named source. In our sample, fewer than four in ten checkable claims could be confirmed. A post with ten statistics and that miss rate is not "mostly accurate"; it is a post with six problems, and readers remember the one they catch. Google evaluates authorship and substance, not the model that produced the draft — what matters is whether the page is worth trusting, and one fabricated figure is enough to lose that trust. It is also the fastest way to produce content that fails every signal of the four-signal slop test. For a related question — whether AI detection is accurate — the answer is also nuanced: see what two studies actually found about AI detectors.
How to catch a made-up statistic in ten minutes
Highlight every number, date and named source in the draft. Search the exact figure alongside the source's name — if the only results are other blogs repeating it, you have a zombie stat. Open the original page, check the number against what it actually measured, and cut anything you cannot verify. The full workflow, including the kill-list of viral stats we keep, is in the fact-check guide.
What we do about it in Skryvo
This test is the reason Skryvo's blog writer runs a fact-check pass by default: each checkable claim in the draft is pulled out, searched, and marked confirmed or unconfirmed, with the source attached. Unconfirmed claims are not hidden; they are flagged so you decide whether to cut, fix or keep them. It does not catch everything (it missed the Litmus figure here), which is why it shows you its sources instead of asking for your trust. You can see how it works on the features page.
Use AI for the draft. Trace every number before you publish. If a statistic cannot be found at its source, it does not go in.
Write content you can publish with confidence
Skryvo generates, fact-checks, and scores your content — try it free.