85% of Hallucinated Citations Survive Into the Published Record

The standard reassurance about AI-fabricated citations is that peer review catches them. Someone reads the manuscript, checks the references, and the invented ones get pulled before publication.

There is now data on whether that happens. It mostly does not.

The study

A preprint led by Zhenyue Zhao, co-authored by arXiv founder Paul Ginsparg, audited 111 million references across 2.5 million papers spanning arXiv, bioRxiv, SSRN, and PubMed Central.

The headline estimate: 146,932 hallucinated citations in 2025 alone, and the authors describe this as conservative.

The more revealing finding came from a different analysis. Because bioRxiv preprints can be traced to their eventual published versions, the team could ask a question nobody had answered directly: what happens to a hallucinated citation between preprint and publication?

85.3% of them survived.

They went into the preprint invented, went through review, and came out the other side still there.

Why that number is the important one

Prevalence figures are contested. Different detection methods produce different estimates, and we have written before about how far apart those estimates can be.

The survival rate sidesteps that problem. It does not depend on knowing the true prevalence. It is a before-and-after comparison on the same set of papers: however many hallucinated citations existed at preprint stage, roughly six in seven were still there after review.

That is a measurement of what review caught, not an estimate of what exists.

The corroborating estimate

Nature‘s news team, working with the screening company Grounded AI, analysed more than 4,000 publications from 2025 across five major publishers — Elsevier, SAGE, Springer Nature, Taylor & Francis, and Wiley.

Their estimate: more than 110,000 publications from 2025 — about 1.65% — contain at least one AI-hallucinated citation.

Worth stating plainly what that figure is: a news-team analysis using a commercial screening tool, not a peer-reviewed count. Treat it as a well-constructed estimate rather than an established fact. It is useful mainly because it was produced independently of the Zhao study and lands in a comparable range.

Why review misses them

Not because reviewers are careless. Because checking references is not what reviewing is.

A reviewer is assessing whether the method is sound, whether the analysis supports the conclusions, whether the contribution is novel. They read the reference list to see whether the relevant literature is engaged with — not to audit whether each entry resolves.

Verifying forty references properly takes one to two hours. Reviewers are unpaid, frequently over-assigned, and evaluated on none of this. Nobody has ever been thanked for catching a fake DOI.

The system was built when fabricating a reference required deliberate fraud. It was never designed to catch fabrication that happens by accident, at scale, from a tool the author trusted.

The Zhao study also found errors clustering in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI writing, and among small and early-career teams — the groups with the least support and the most pressure to publish.

What follows from this

Publication is not verification. A citation appearing in a peer-reviewed paper tells you a reviewer did not object to it. On this evidence, that is a weak signal about whether it exists.

Citing from other papers’ reference lists carries risk. The habit of pulling a citation from a published paper’s bibliography without opening the source now has a measurable failure rate attached to it.

The check has to happen before submission. Not because reviewers are failing, but because they were never doing this job. Our verification checklist is the version that takes two minutes per reference.

There is a structural fix available and it is not complicated: automated DOI resolution at submission, before a manuscript reaches a reviewer. The data is public, the check is mechanical, and it would catch the majority of these before a human ever sees the paper.

Until publishers implement that, the burden sits with authors — which is to say, with you.


Sources

  • Zhao et al. — arXiv:2605.07723 — audit of 111 million references across four preprint and publication corpora. Preprint, not peer reviewed.
  • Nature — investigation with Grounded AI into hallucinated citations in the published literature
  • Grounded AI — methodology behind the Nature analysis

Figures are as reported by the studies cited; we have not independently replicated them. The Zhao study is a preprint and has not completed peer review. Research-based rather than hands-on — see our Methodology page.