The Retraction That Keeps Getting Cited
The way we caught the paper mills is wild, and it is also why I reckon we are not fixing the problem. One syndicated trawl found papers solemnly defining “kidney disappointment” as a disease, as in “Kidney disappointment is an ailment in which the kidneys no longer function.”1 It reported about 189 hits for that exact phrase, many looking like mistranslations of kidney failure.1 It is funny until you see what it is a fingerprint for.
That fingerprint has a name. Sleuths have found thousands of “tortured phrases” littered throughout reputable journals.2 They typically result from using paraphrasing tools to evade plagiarism-detection software when stealing someone else’s text.2 Think “counterfeit consciousness” for artificial intelligence or “colossal information” for big data, the kind of synonym substitution a thesaurus does when it has no idea what the words mean. Guillaume Cabanac built a detector that now combs through 130 million publications every week and has been instrumental in over 1,000 retractions.3 Even its narrow filter for papers with at least five tortured phrases each has turned up nearly 19,000 papers.3 The curve is going the wrong way: using the same criteria there were 416 such cases in 2018, 638 in 2019, and 1,135 in 2020.4 Over the past decade, that cottage trick has been industrialised into proper paper mills selling bogus research at scale.5 The clean-up is now priced in. Wiley told investors its $18 million drop in research publishing revenue was mainly due to the Hindawi disruption.6 After that it retracted more than 11,300 Hindawi papers in two years.6 When researchers re-examined why papers get retracted, fraud accounted for over 43 percent.7
It is tempting to say detection is working, so correction will follow. In one sense it is. Springer Nature now says if a number of non-standard phrases are identified by its tool, the submission will be withdrawn.8 Cabanac’s own nightly trawl of 120 million records in Dimensions keeps ferreting out tortured expressions before they calcify. The 2026 World Conference on Research Integrity pulled more than 700 people from 57 countries to talk about AI and research security, and its workshops were blunt about slow investigations, confidentiality barriers and inconsistent institutional responses.9 Even outside biomedicine the strain shows: at NeurIPS, a detector found 100 hallucinations in more than 51 accepted papers.10 While submissions more than tripled and the average mistakes per paper rose from 3.8 to 5.9, you can see why reviewers are swamped. I get the pushback, and I have heard it from lab friends too: this is a peer review capacity problem, an AI writing problem, a publish-or-perish problem, and better screening will sort it. Some even argue the humaniser tools that strip AI tells from text are just helping clarity, not cheating. I do not buy that as an answer to the citation problem, because the detectors finding more does not mean the literature is healing faster.
Here is the bit I keep coming back to. More than 764,000 articles cite retracted works, and about 5,000 of those have at least five retracted references in their bibliographies.3 Separately, a CNRS audit flagged some 700 oncology articles published between 2014 and 2018 that included errors in genetic sequences, and are nonetheless cited more than 20,000 times in other scientific articles.4 Retraction is not erasure, it is a flag on a record that the citation graph ignores. That matters most where it hurts: fake papers are slowing research that has helped millions with lifesaving medicine, from cancer to COVID, and analysts’ data shows cancer and medicine are particularly hard-hit.5 The mills have got good at producing plausible nonsense, the detectors have got good at finding the tells, but the system for stopping the downstream use has not moved. Until a retraction or flag reliably stops the next paper citing the flawed one, automated screening is just counting the spill while the floor stays wet.
Sources
- 1 Google Scholar search reveals ‘kidney disappointment’ in research papers — Research papers using “kidney disappointment… — zeli.app
- 2 @aix1 on Research papers using “kidney disappointment” instead of “kidney failure” — ycombinator.com
- 3 Problematic Paper Screener: Trawling for fraud in the scientific literature — theconversation.com
- 4 Guillaume Cabanac tracks fake science | CNRS News — cnrs.fr
- 5 Bogus research is undermining good science, slowing lifesaving research - Ars Technica — arstechnica.com
- 6 Wiley shuts 19 scholarly journals amid AI paper mill problem — theregister.com
- 7 Research fraud exploded over the last decade - Ars Technica — arstechnica.com
- 8 Springer Nature expands its portfolio of research integrity tools to detect non-standard phrases | Springer Nature Group | Springer Nature — springernature.com
- 9 Reporting from the 2026 World Conference on Research Integrity – Science Integrity Digest — scienceintegritydigest.com
- 10 AI conference’s papers contaminated by AI hallucinations — theregister.com
How this was made
- 01-research z-ai/glm-5.2 $0.075
- 03-annotate z-ai/glm-5.2 $0.031
- 04-nominate deepseek/deepseek-v4-pro $0.018
- 05-select google/gemini-3.7-flash $0.003
- 06-write meta/muse-spark-1.2 $0.065
- 08-visualise anthropic/claude-sonnet-5 $0.039
total $0.232