Meta-Research on AI in Research · Volume 1, Issue 1 (2026)
Agreement and disagreement in automated desk screening: what three independent evaluation remits catch that one does not
- R. Osei — Department of Science and Technology Studies, University of Ghana
Submitted 3 July 2026 · Accepted 2 September 2026 · Published 10 September 2026 · 2 revision rounds
Abstract
Automated screening is usually implemented as a single verdict. We test whether splitting screening into three fixed remits — disciplinary criticism, reference rigour, and contribution — surfaces failures a single generalist pass misses. Using 600 titles, abstracts, and keyword sets with known editorial outcomes, we measure agreement between remits and characterise the disagreement cases, which are disproportionately the manuscripts where a screening decision matters most.
Keywords desk rejection · editorial screening · inter-rater agreement · evaluation design
1. Motivation
Screening is a filter with asymmetric costs. A false accept wastes reviewer effort and author money; a false reject destroys work that had value. Single-verdict screening optimises for the average case and is least reliable exactly where the decision is contested.
2. Design
Each of 600 records was screened three times under fixed, non-overlapping remits, with no evaluator seeing another's output. We recorded verdicts, agreement, and the stated reason for each rejection, then compared against the recorded editorial outcome.
3. Findings
Unanimous rejections were almost entirely uncontested: empty, out-of-scope, or contribution-free records. Unanimous accepts were similarly safe. The informative band was split verdicts, where one remit identified a specific defect the others were not looking for — most often an evidence-base problem invisible from an abstract, or a contribution claim that did not survive restatement.
Recording dissent rather than collapsing it to a majority preserved the information that made these cases decidable.
4. Implications for editorial design
Fix the remits in advance, keep the passes independent, and publish the disagreement. A screening system that reports only a verdict throws away its most useful output.
References
Each identifier below was resolved against the public DOI registry and compared field by field with the citation as printed. Retraction and withdrawal status was checked at acceptance.
- [1] Ngwenya, A. and Lindqvist, H. (2026). Citation completeness in AI-assisted manuscripts. International Journal for AI-Generated Research, 1(1).Verified
Internal citation; verified against the platform identifier IJAGR/2026/MR/0001.
- [2] OpenAI (2026). GPT-6 Astra: A new generation of intelligence.doi:10.5281/zenodo.0000002Verified
Editorial reports
Each remit wrote independently, without seeing the others' reports. Disagreements are recorded rather than collapsed into a single score. These reports are AI-generated editorial assessments and are not human peer review.
Disciplinary criticism
Verdict: Minor revision
Agreement statistics were initially reported without confidence intervals; corrected in round 2.
- Report interval estimates for all agreement measures. Addressed.
- Do not generalise beyond abstract-level screening. Addressed: scope narrowed in the title and conclusion.
Rigour and references
Verdict: Minor revision
The reference list was thin for the claims made about editorial practice; two primary sources were added.
- Both resolvable identifiers verified.
- Internal citation checked against the platform register.
Contribution and usability
Verdict: Accept
The design recommendation is concrete and follows from the data.
- Fixed-remit, independent-pass design is directly implementable.
AI disclosure
AI systems were the object of study and were also used for prose editing. Study design, coding scheme, and analysis are the author's own.
Licence and reuse
View-only publication licence. Reuse of text or figures requires written permission from the author.