This publication has spent most of its existence on verification. How to check citations. Why disclosure policies fail. What detection catches and what it misses.
There is a reasonable argument that all of it addresses the wrong layer.
The case being made
Chemical & Engineering News put it directly in its July 2026 examination of AI in scientific publishing: detectors may never beat fraud, and changing how science rewards its researchers might be the only fix that does.
The reasoning runs like this. Generative AI did not create the problem. It made an existing one cheap.
Researchers are evaluated primarily on how many papers they publish and how prestigious the journals are. Tenure depends on it. So does funding, promotion, salary, institutional ranking — and, for international researchers, immigration status. That system predates AI by decades and has been criticised for just as long.
Ivan Oransky, co-founder of Retraction Watch, puts it bluntly: we do not need generative AI to make a dog’s breakfast out of the scientific literature.
What AI changed is the cost. Fabricating a paper used to take effort. Now it does not. When you make something cheap that was already incentivised, you get more of it.
Why detection cannot win
The detection industry is unusually candid about this. Its own developers acknowledge that the race against improving models is one they may not win.
That is structural, not pessimism. Detectors train on the output of current models. Models improve. The gap closes, then reopens somewhere else. Meanwhile the cost asymmetry favours the generator: producing a fake paper is fast and cheap, while detecting it reliably at scale is neither.
We have already seen how wide the uncertainty is. Two methods analysing overlapping samples of ICLR 2026 peer reviews returned figures four times apart — 21% versus 5%. When the measurement disagrees with itself by that margin, enforcement built on it inherits the disagreement.
Disclosure has the same weakness from a different angle. It depends on author honesty in a system where the incentive runs toward concealment. Analysis of more than 5 million papers found no measurable difference in AI-assisted writing between journals with disclosure policies and those without.
The scale of what is being defended against
Paper mills — commercial operations that fabricate research and sell authorship — are the clearest illustration of incentives operating as designed.
A study published in August 2025 reported the number of paper-mill publications doubling roughly every 18 months, far outpacing growth in legitimate output. A COPE and STM assessment estimated around 2% of submitted articles originate from paper mills, with some analyses suggesting rates as high as 20% in particular journals and disciplines.
Paper mills are not a technology failure. They are a market. They exist because there is demand, and there is demand because publication count determines careers.
What changing the incentive would look like
The proposal most often raised is narrower than it sounds.
Guillaume Cabanac, the computer scientist behind the Problematic Paper Screener, describes his version: if he had a magic wand, researchers would list only their top five articles, with evaluation focused on quality rather than quantity. He notes this is not hypothetical — in France, hiring and promotion committees are already asked to focus on a short selection of works.
The mechanism is simple. If only your best five papers count, a sixth mediocre paper is worth nothing. Publishing volume stops being a strategy. The market for fabricated authorship contracts because the thing it sells no longer converts into career advantage.
No detector required.
Where the argument is weaker than it sounds
Worth pushing back on, because “fix the incentives” can function as a way of deferring everything.
It is slow. Evaluation norms are set by thousands of independent institutions, national systems, and funders. France moving does not move Germany. Reform on that scale takes decades, and the fabricated-citation rate is rising now.
Quality assessment has its own failure modes. Judging five papers on merit requires people with time, expertise, and no conflicts. That is a scarce combination, and “quality” assessed badly can favour the well-connected more reliably than counting does.
It does not cover honest error. A researcher who trusts an AI-suggested citation and does not check it has not gamed any incentive. Removing publication pressure entirely would not prevent that. Verification remains necessary regardless of what the reward structure looks like.
Both layers, honestly labelled
The useful conclusion is not that verification is pointless. It is that verification and incentive reform address different problems, and conflating them produces bad expectations.
Verification handles error. The checklist catches a fabricated citation you accepted in good faith. It protects your own work, it is entirely within your control, and it takes minutes. That is real and worth doing.
Incentive reform handles fraud. Nobody running a paper mill is going to be stopped by a verification checklist. That is a market responding to demand, and it shrinks only when demand does.
Most of what this site has published sits in the first category. That is the layer an individual researcher can actually act on, which is why it deserves the attention.
But it is worth saying plainly: careful individual verification will not fix the scientific literature. It will keep your own contribution to it clean. Those are different achievements, and the second one is the honest ceiling on everything else here.
Sources
- Chemical & Engineering News — Can we stop AI from flooding scientific journals? Includes quotes from Ivan Oransky and Guillaume Cabanac.
- How to fight fake papers — review of paper mill drivers and information sources
- Frontiers in Medicine — Integrity at stake: confronting publish or perish, including COPE/STM paper mill submission estimates
Quotes are as reported by C&EN. Paper mill prevalence estimates vary widely by methodology and discipline — treat the ranges as indicative rather than precise. Research-based rather than hands-on — see our Methodology page.