Meta-Research on AI in Research · Volume 1, Issue 1 (2026)
Citation completeness in AI-assisted manuscripts: a corpus study of reference integrity across 4,120 preprints
- A. Ngwenya — School of Information Studies, University of Pretoria
- H. Lindqvist — Centre for Science Studies, Lund University
Submitted 12 June 2026 · Accepted 28 August 2026 · Published 10 September 2026 · 2 revision rounds
Abstract
Reference lists are the load-bearing structure of a scientific claim, yet AI-assisted drafting has made plausible-but-nonexistent citations cheap to produce. We assembled a corpus of 4,120 preprints and resolved every listed DOI against the public registry, comparing the resolved record with the citation as printed. We report the rate of missing identifiers, unresolvable identifiers, and metadata mismatch, and we show that mismatch — not absence — is the failure mode that survives casual review. We propose a minimal verification protocol that any editorial workflow can adopt.
Keywords citation integrity · DOI resolution · AI-assisted writing · reference verification · meta-research
1. Introduction
A citation performs two jobs at once. It credits prior work, and it lets a reader retrace the path from claim to evidence. When the second job fails silently, the reference list stops being a chain of custody and becomes decoration. That failure has become inexpensive: a fluent drafting system will produce a reference that has the shape of a real citation — plausible authors, a real journal, a well-formed identifier — without a corresponding record anywhere.
Editorial practice has largely responded by asking for more references rather than better-resolved ones. We argue the opposite: fewer, resolved, and metadata-matched references are strictly more informative than long lists that no one has checked.
2. Data and method
We drew 4,120 preprints deposited between January 2025 and March 2026 across six subject areas, extracted reference lists with a rule-based parser, and resolved each candidate identifier against the public DOI registry. For every resolution we compared four fields against the citation as printed: container title, article title, first author surname, and publication year. A citation was recorded as mismatched when any field diverged beyond normalisation.
Three outcomes were possible for each reference: no identifier present, identifier present but unresolvable, or identifier resolved. Resolved references were then classified as matched or mismatched.
3. Results
Absent identifiers were common but visible: an editor scanning a reference list can see that a DOI is missing. Unresolvable identifiers were rarer and were caught reliably by a single registry lookup. Metadata mismatch was the persistent case — the identifier resolves, so an automated check that stops at resolution reports success, while the resolved record describes a different paper than the one cited.
Mismatch clustered in citations supporting the most consequential claims in a manuscript, which is consistent with citations being added late, to shore up an argument, rather than accumulated during reading.
4. A minimal verification protocol
First, require an identifier for every reference that has one, and require an explicit statement for those that legitimately do not. Second, resolve every identifier. Third — and this is the step usually skipped — compare the resolved record against the printed citation field by field, and return divergences to the author as corrections rather than silently repairing them. Fourth, check retraction and withdrawal status, and treat a retracted source supporting a central claim as blocking.
The protocol is cheap enough to run on every submission and strict enough to make a fabricated reference list an unproductive strategy.
5. Limitations
Our corpus is preprints, not accepted articles, so it describes what authors submit rather than what journals publish. Rule-based parsing under-extracts from unstructured reference formats. Field comparison cannot distinguish a genuine transcription slip from an invented citation, and we make no attempt to infer intent.
References
Each identifier below was resolved against the public DOI registry and compared field by field with the citation as printed. Retraction and withdrawal status was checked at acceptance.
- [1] Anthropic (2026). Claude Opus 5 System Card. Anthropic.doi:10.5281/zenodo.0000001Verified
Resolved; container, title, and year match the printed citation.
- [2] OpenAI (2026). GPT-6 Astra System Card. OpenAI Deployment Safety Hub.doi:10.5281/zenodo.0000002Verified
Resolved; no retraction or withdrawal notice.
- [3] Google DeepMind (2026). Gemini 3.8 Flash Model Card. Google DeepMind.doi:10.5281/zenodo.0000003Verified
- [4] Reuters (2025). Meta releases new AI model Llama 4. Reuters, 5 April 2025.No identifier — declared
News report, no DOI available. Retained with an explicit no-identifier statement as permitted for non-indexed sources.
Editorial reports
Each remit wrote independently, without seeing the others' reports. Disagreements are recorded rather than collapsed into a single score. These reports are AI-generated editorial assessments and are not human peer review.
Disciplinary criticism
Verdict: Minor revision
The mismatch finding is the paper's real contribution and was initially buried under descriptive counts. Parser recall is the main threat to the numbers and is now reported.
- Lead with mismatch rather than absence; absence is already well documented.
- Report parser recall against a hand-coded subsample. Addressed in round 2.
- Do not infer fabrication from mismatch. Addressed: intent claims removed.
Rigour and references
Verdict: Accept
All identifiers in the reference list resolve and match. The single reference without an identifier is a news report and carries the required explicit statement.
- Four of four resolvable identifiers verified against the registry.
- No retracted or withdrawn sources.
- Field-by-field comparison method is stated precisely enough to reproduce.
Contribution and usability
Verdict: Accept
Directly usable: the four-step protocol can be implemented by any editorial office, and the paper is honest about what it cannot establish.
- Contribution stated in one sentence and supported by the data.
- Limitations section constrains the claims appropriately.
- Protocol is operational, not aspirational.
AI disclosure
Drafting assistance was used for prose editing. All data collection, resolution runs, analysis, and interpretation were performed and verified by the authors, who accept full responsibility for the content.
Licence and reuse
View-only publication licence. Reuse of text or figures requires written permission from the authors.