The infrastructure of integrity: What does it take to trust a published research output?

Research integrity is supported by a chain of interconnected infrastructure. Publishers, funders, and researchers depend on this chain every day, yet many of the critical links remain invisible unless they break. This analysis traces the infrastructure required to support research integrity in the context of increasing challenges created by generative AI. These infrastructures are more critical than ever, and many are under growing strain.

Generative AI is prompting a research integrity crisis

Preprint Flooding

LLM-generated manuscripts submitted en masse to preprint servers are overwhelming review systems. Fabricated data, hallucinated citations, and plausible-sounding methodology can pass automated checks and cursory human screening (Singh, 2025).

Identity Theft & Ghost Authors

Paper mills and fraudulent or low-quality journals falsely attribute authorship to real researchers to lend credibility (Grove, 2023).

AI-Powered Paper Mills

Organised networks now use generative AI to produce hundreds of superficially coherent papers monthly, defeating manual peer review through sheer volume and automated reviewer impersonation (Richardson et al., 2025).

Citation Chain Corruption

AI systems hallucinate references that appear in real papers, creating citation chains that lead to non-existent or misrepresented sources, potentially leading to large-scale pollution of the scholarly record (Zhao et al., 2026).

Open Infrastructure interventions bolster trustworthy knowledge

01

Identify

Who is the researcher?
ORCID
AI risk
AI-generated papers steal the identities of real researchers to lend credibility to fabricated or low-quality publications.
OI intervention
Persistent, globally unique researcher identifiers create a trustworthy link between an individual and their work. Open identity infrastructure makes impersonation and ghost authorship detectable. An ORCID — a free, unique identifier that follows a researcher for their entire career regardless of name changes, or institutional moves — can prevent fraudulent authorship claims. Authentication via ORCID when authors submit data, manuscripts, or other outputs, provides a layer of identity validation.
02

Affiliate

Where does the work come from?
ROR ImpactU
AI risk
Fabricated affiliations allow paper-mill output to look like it came from research-intensive and prestigious institutions.
OI intervention
ROR (Research Organization Registry) assigns a persistent, openly-licensed identifier to over 116,000 institutions worldwide. ROR is structurally embedded in Crossref DOI metadata, DataCite DOI metadata, and ORCID, making it much more difficult to assert false affiliations when these persistent identifiers are used in concert. Current Research Information Systems (CRIS) like ImpactU make affiliations visible by tracking verified institutional output. This permits cross-checking of affiliation claims against an institution’s actual, indexed publication record.
03

Connect

What other outputs support this work?
Dryad DMP Tool DataCite
AI risk
It is easier than ever to fabricate or alter datasets using generative AI tools. Falsified data can be used to support misleading, untrue, or exaggerated claims in published research.
OI intervention
Data Management Plans (DMPs) enforce pre-registration of methodology before results are known. DMPTool, for example, provides a click-through wizard for creating a DMP that complies with funder requirements, making after-the-fact alterations easier to spot. Open data repositories and platforms make datasets available for validation by the research community. At Dryad, for example, submissions are self-deposited, then vetted by Dryad curators and given a citable DOI. The published dataset becomes persistently available for peer review and validation. As the primary DOI registration agency specializing in research data and software, DataCite provides persistent identifiers that link datasets to research articles and other outputs, building an interconnected bundle of evidence and analysis that can be independently validated.
04

Propagate

Metadata flows into the scholarly record
Crossref OpenAlex
AI risk
Papers (whether AI-generated or assisted) contain hallucinated citations.
OI intervention
Persistent Identifiers (PIDs) such as DOIs and handles create easy, timestamped metadata about citations that AI tools could use and research integrity services can use to verify citations. Crossref’s 2026 public data file contains nearly 180 million structured, machine-readable DOI records, providing an authoritative source of truth on the published record. OpenAlex, a free, open-source index of hundreds of millions of scholarly works, provides another portal to verifying citations.

A chain of interconnected infrastructure

01
Identify
Who is the researcher?
ORCID
02
Affiliate
Where does the work come from?
ROR ImpactU
03
Connect
What other outputs support this work?
Dryad DMP Tool DataCite
04
Propagate
Metadata flows into the scholarly record
Crossref OpenAlex

What needs to happen

Generative AI lowers the cost of producing poor-quality and fraudulent research by orders of magnitude. Research integrity, trust in science, and the foundations of how we communicate knowledge are at risk. As this analysis shows, we already have the infrastructure we need to face this risk — but only if it’s adequately funded, adopted, and maintained. Free to use does not mean free to run.

Funders, institutions, and libraries must sustain tools that help to ensure research integrity. Sustained investment ensures the shared, open foundations that the entire research ecosystem depends on.

Build PIDs into policy.

Top-down requirements from public and private funders can help to build adoption. Mandate ORCID iDs for PIs and co-investigators, ROR IDs for all named institutions, and DOIs for funded datasets and outputs.

Fund the infrastructure directly, not just the research it supports.

Infrastructures like ORCID, ROR, DataCite, Crossref, and Dryad are largely nonprofit, community governed, and chronically underfunded relative to how much the research ecosystem depends on them. Fund institutional memberships rather than passing off costs to researchers to absorb individually.

Treat open infrastructure maintenance as a budget line, not a one-time subscription.

Sustained, multi-year commitments allow open infrastructure to plan capacity to keep pace with AI-driven changes.

Scholarly communications relies on shared open infrastructure to provide critical integrity functions. This infrastructure works because it’s shared, which means its cost has to be shared too. Sustained, coordinated investment from funders, institutions, and publishers is what keeps identity, affiliation, data, and citation infrastructure ahead of AI-driven risks that impact the entire knowledge ecosystem.

Works Cited

  1. Grove, Jack. (2023, February 2). Identity theft victims back legal action against journals. Times Higher Education (THE). https://www.timeshighereducation.com/news/identity-theft-victims-back-legal-action-against-journals
  2. Richardson, Reese A. K., Hong, Spencer S., Byrne, Jennifer A., Stoeger, Thomas, & Amaral, Luís A. Nunes. (2025, August 12). The entities enabling scientific fraud at scale are large, resilient, and growing rapidly. Proceedings of the National Academy of Sciences (PNAS). https://pubmed.ncbi.nlm.nih.gov/40758886/
  3. Singh, Ranjit. (2025, October 30). On arXiv, an Influx of AI Slop Pits Surface Against Substance. Data & Society. https://datasociety.net/research-library/on-arxiv-an-influx-of-ai-slop-pits-surface-against-substance/
  4. Zhao, Zhenyue, Wang, Yihe, Stuart, Toby, De Vaan, Mathijs, Ginsparg, Paul, & Yin, Yian. (2026, May 8). LLM hallucinations in the wild: Large-scale evidence from non-existent citations. arXiv. https://arxiv.org/pdf/2605.07723

Feedback