The Counterfeit Citation
The academic citation is the atomic unit of scientific trust. Every claim in a published paper rests on a chain of references — verifiable links to prior work that let readers check the foundations. That chain is now being counterfeited at industrial scale, and the institutions built on trust are scrambling to respond.
A peer-reviewed study published in The Lancet by a Columbia University team analyzed 2.5 million papers from PubMed Central containing over 125 million references. The findings are staggering: fabricated citations rose roughly 12-fold in three years. In 2023, one in 2,828 papers contained at least one fake reference. By 2025, it was one in 458. By early 2026, one in 277. The total: over 4,000 fabricated citations across 2,800 papers. Lead author Maxim Topaz described the real-world stakes bluntly: “A medical professional or clinical guideline developer has no way of knowing that the evidence they are relying on does not exist.” One paper he reviewed had 18 out of 30 fake references — and some of those phantom citations were already being cited by other papers and appearing in systematic reviews that inform clinical care.
The Anatomy of a Fake
What makes this problem insidious is how convincing the fakes have become. Early AI hallucinated citations were easy to spot — garbled titles, wrong fields, obvious nonsense. The current generation is something else entirely. A Nature investigation conducted with screening company Grounded AI found that modern hallucinated references assemble fragments of genuine publications into what CEO Joe Shockman calls “Frankenstein” citations: titles matching an author’s research area, co-authors who have published together before, volume and page numbers that sync with the journal and publication date. AI also hallucinates DOIs — the unique digital identifiers that are supposed to be the bedrock of citation verification. Kathryn Weber-Boer, director of scientometrics at Digital Science, confirmed: “AI also hallucinates DOIs, both in references that are otherwise genuine as well as in fabricated ones.”
The contamination is showing up everywhere. At ICLR 2026 — one of the three premier AI conferences — Mark Russinovich’s RefChecker tool found that one in 29 accepted papers contained at least one likely hallucinated reference. GPTZero scanned 4,841 NeurIPS 2025 papers and found at least 100 hallucinated citations across 51 accepted papers. A systematic analysis of four major HPC conferences found that every 2025 proceeding contained mysterious citations — affecting 2–6% of published papers — while the 2021 proceedings contained none. One citation even included a URL with utm_source=chatgpt.com, a smoking gun the authors hadn’t bothered to remove. A Google Scholar search found nearly 17,000 articles containing that same tracking parameter in their references.
The Nuclear Option
Into this crisis steps Thomas Dietterich, chair of arXiv’s computer-science section, with what amounts to the nuclear option: a one-year ban from arXiv for any submission containing hallucinated references. After the ban, authors must have subsequent submissions accepted at a reputable peer-reviewed venue before they can submit to arXiv again. The policy sent shockwaves through the research community — not because it’s unreasonable, but because of what it signals about the scale of the problem arXiv is seeing behind the curtain.
The HN community split predictably. Defenders called it “a very modest raising of the bar” that comes at “zero cost to honest researchers.” Critics pushed back on the punitive framing: “It’s just one darn hallucinated citation… It doesn’t account for the substance or quality of their work at all.” One commenter raised the scenario of a researcher asking an AI to find a citation they know exists, only for the AI to silently insert an incorrect one without them noticing. Dietterich’s response cut to the core: “Our Code of Conduct states that by signing your name as an author of a paper, each author takes full responsibility for all its contents, irrespective of how the contents were generated.” The question isn’t whether you used AI. The question is whether you checked its work.
This is the right framing, and also the uncomfortable one. The entire peer-review system was designed for a world where humans wrote their own bibliographies. Reviewers don’t verify references — they can’t, at the scale and pace of modern publishing. The HPC study’s authors noted that no author in their dataset acknowledged using AI to generate citations, despite conference policies requiring disclosure. Current disclosure policies are, in their words, “insufficient.”
The Deeper Problem
What makes this different from previous academic integrity scandals is the mechanism. This isn’t plagiarism. It isn’t data fabrication by a rogue researcher. It’s a systemic contamination vector that exploits the trust architecture of science itself. The citation graph — the web of references connecting papers to their foundations — is the closest thing academia has to a knowledge verification system. When a paper cites a source, it’s making a claim: “this prior work exists and supports my argument.” When that claim is false, and the false citation gets cited by other papers, which get cited by others, the corruption compounds. It’s a citation laundering operation run by machines that don’t know they’re laundering.
The Lancet study found that some fabricated citations were already appearing in systematic reviews — the gold-standard evidence summaries that inform clinical guidelines. In medicine, a phantom citation isn’t an abstract integrity problem. It’s a patient safety issue. A doctor following a clinical guideline that references a study that doesn’t exist is making decisions based on evidence that was never produced.
The Detection Arms Race
The response is bifurcating into two tracks: punishment and detection. arXiv is going the punishment route. Companies like Grounded AI are building automated screening tools. The Columbia team recommends that journals implement automated reference verification before peer review, that indexing services add integrity metadata to article records, and that publishers retroactively screen existing publications. The RefChecker tool that found the ICLR problems is open-source — anyone can run it.
But detection is inherently reactive. The fundamental problem is that LLMs produce confident, plausible, well-formatted nonsense, and the humans using them are either too busy, too trusting, or too careless to verify. As one HN commenter put it: “If you use AI correctly, nobody should be able to tell that it was used at all.” The flip side is that if you use AI carelessly, nobody can tell until someone checks every reference — and almost nobody does.
The arXiv ban is a signal flare. It tells the research community that the preprint server — the backbone of rapid scientific communication, particularly in AI and physics — is seeing enough hallucinated references to warrant draconian enforcement. Whether a one-year ban is proportionate is a legitimate debate. Whether the problem demands aggressive action is not. The scientific literature is accumulating phantom knowledge at an accelerating rate, and every day it goes unaddressed, the citation graph becomes a little less trustworthy.
Sources
- AI Blamed For Rise In Fabricated Citations Found In Recent Research Papers — Forbes
- Fabricated references in research papers — The Lancet
- The Case of the Mysterious Citations — arXiv
- Hallucinated Citations Are Polluting the Scientific Literature — Nature/Longreads
- ICLR 2026 RefChecker findings — Mark Russinovich on LinkedIn
- New arXiv policy: 1-year ban for hallucinated references — Hacker News
- Over fifty new hallucinations in ICLR 2026 submissions — Hacker News