Nearly every major journal now has a policy on AI use. Disclose it, or don’t use it. The rules exist, they are published, and they are easy to find.
A study published in PNAS in March 2026 suggests they are not doing much.
What the study found
Yongyuan He and Yi Bu of Peking University analysed more than 5.2 million papers across 5,114 journals. They sorted those journals by policy: 3,556 require disclosure of AI use, 1,529 say nothing about AI at all, 27 prohibit it outright, and 2 have an explicitly open policy.
Then they measured how much AI-assisted writing actually appeared in each group over time.
The growth curves were, in the paper’s own words, statistically indistinguishable. There was no significant difference between journals with policies and journals without them. Bu summarised it as having “no statistically measurable effect.”
The disclosure numbers point the same direction. In the full-text sample, 76 out of 75,172 post-2023 papers disclosed AI use — roughly one in a thousand. Given what the same data suggests about actual AI usage rates, that gap is the story.
Why a policy alone doesn’t change behaviour
A disclosure requirement asks authors to volunteer information that can only hurt them. There is no upside to disclosing, no reliable way for the journal to detect non-disclosure, and no consequence for staying quiet.
That is not a moral failing peculiar to researchers. It is what happens to any rule with no detection mechanism and no penalty. The policy documents the norm; it does not enforce it.
arXiv tried enforcement instead
In May 2026, arXiv took a different approach. Thomas Dietterich, who chairs the computer science section, announced a one-strike rule: submit a paper with incontrovertible evidence of unchecked AI generation, and you are banned from submitting for a year. After that, future submissions must first be accepted at a peer-reviewed venue.
“Incontrovertible evidence” means what it sounds like. Hallucinated references. Leftover chatbot meta-commentary sitting in the manuscript — the now-infamous “here is a 200-word summary; would you like me to suggest alterations?” variety.
arXiv separately moved to reject most computer science review and survey papers that cannot show prior peer review. Coverage of the change ran in 404 Media and Inside Higher Ed, where the reaction was summarised as welcome but unenforceable.
That criticism has weight. The rule only catches the careless — authors who left the evidence visible. A researcher who checks their references and deletes the chatbot artefacts is invisible to it. Enforcement bites at the bottom of the distribution.
It is still a meaningfully different instrument from a disclosure request. One asks. The other costs something.
The detection problem underneath all of this
Both approaches assume someone can tell when AI was involved. That assumption is shakier than it looks.
Consider the analysis of peer reviews submitted to ICLR 2026. Pangram Labs examined roughly 75,800 reviews and classified 21% as fully AI-generated. A competing method, Sem-Detect, ran on an overlapping sample and returned 5%.
Same reviews. Two detection methods. A fourfold gap. Pangram also sells AI-detection tooling, which is worth stating plainly when weighing the higher figure.
If the field cannot agree on how much AI-generated text is in a fixed corpus, enforcement built on detection inherits that uncertainty.
Where this leaves individual researchers
Two things follow, and they point in the same direction.
Disclosure norms are unsettled, so read your target journal’s actual policy rather than assuming. Requirements vary from silence to outright prohibition, and “everyone does it” is not a defence if your journal is one of the 27 that prohibits it.
The thing that gets people caught is not disclosure — it is artefacts. Fabricated references and leftover chatbot text are what triggered arXiv’s rule. Those are verification failures, not policy failures, and they are entirely within your control. Our guide to verifying AI-generated citations covers the checks that catch them.
Policy is downstream of detection, and detection is contested. Verification is the part that works regardless of which way the rules settle.
Sources
- He, Y. & Bu, Y. — Academic journals’ AI policies fail to curb the surge in AI-assisted academic writing, PNAS (2026), DOI 10.1073/pnas.2526734123
- PNAS Science Sessions — interview with the study authors
- 404 Media — arXiv changes rules after AI-generated submissions
- Inside Higher Ed — reaction to the arXiv ban
- Pangram Labs — ICLR 2026 review analysis
- Sem-Detect — competing detection method, arXiv:2605.21713
Research-based rather than hands-on. See our Methodology page for what that distinction means and how we evaluate.