The infrastructure of integrity: What does it take to trust a published research output?
Generative AI is prompting a research integrity crisis
Preprint Flooding
LLM-generated manuscripts submitted en masse to preprint servers are overwhelming review systems. Fabricated data, hallucinated citations, and plausible-sounding methodology can pass automated checks and cursory human screening (Singh, 2025).
Identity Theft & Ghost Authors
Paper mills and fraudulent or low-quality journals falsely attribute authorship to real researchers to lend credibility (Grove, 2023).
AI-Powered Paper Mills
Organised networks now use generative AI to produce hundreds of superficially coherent papers monthly, defeating manual peer review through sheer volume and automated reviewer impersonation (Richardson et al., 2025).
Citation Chain Corruption
AI systems hallucinate references that appear in real papers, creating citation chains that lead to non-existent or misrepresented sources, potentially leading to large-scale pollution of the scholarly record (Zhao et al., 2026).
Open Infrastructure interventions bolster trustworthy knowledge
Identify
Affiliate
Connect
Propagate
A chain of interconnected infrastructure
What needs to happen
Generative AI lowers the cost of producing poor-quality and fraudulent research by orders of magnitude. Research integrity, trust in science, and the foundations of how we communicate knowledge are at risk. As this analysis shows, we already have the infrastructure we need to face this risk — but only if it’s adequately funded, adopted, and maintained. Free to use does not mean free to run.
Funders, institutions, and libraries must sustain tools that help to ensure research integrity. Sustained investment ensures the shared, open foundations that the entire research ecosystem depends on.
Top-down requirements from public and private funders can help to build adoption. Mandate ORCID iDs for PIs and co-investigators, ROR IDs for all named institutions, and DOIs for funded datasets and outputs.
Infrastructures like ORCID, ROR, DataCite, Crossref, and Dryad are largely nonprofit, community governed, and chronically underfunded relative to how much the research ecosystem depends on them. Fund institutional memberships rather than passing off costs to researchers to absorb individually.
Sustained, multi-year commitments allow open infrastructure to plan capacity to keep pace with AI-driven changes.
Works Cited
- Grove, Jack. (2023, February 2). Identity theft victims back legal action against journals. Times Higher Education (THE). https://www.timeshighereducation.com/news/identity-theft-victims-back-legal-action-against-journals
- Richardson, Reese A. K., Hong, Spencer S., Byrne, Jennifer A., Stoeger, Thomas, & Amaral, Luís A. Nunes. (2025, August 12). The entities enabling scientific fraud at scale are large, resilient, and growing rapidly. Proceedings of the National Academy of Sciences (PNAS). https://pubmed.ncbi.nlm.nih.gov/40758886/
- Singh, Ranjit. (2025, October 30). On arXiv, an Influx of AI Slop Pits Surface Against Substance. Data & Society. https://datasociety.net/research-library/on-arxiv-an-influx-of-ai-slop-pits-surface-against-substance/
- Zhao, Zhenyue, Wang, Yihe, Stuart, Toby, De Vaan, Mathijs, Ginsparg, Paul, & Yin, Yian. (2026, May 8). LLM hallucinations in the wild: Large-scale evidence from non-existent citations. arXiv. https://arxiv.org/pdf/2605.07723