Local AI pre-screening for human triple-blind peer review in health sciences
Rodrigo Martins Boos * iD
Publicado em 17/07/2026 · Ahead of print · DOI: 10.5281/zenodo.21365017 · 18 visualizações
Resumo
Academic peer review is under mounting strain: NeurIPS 2025 (Conference on Neural Information Processing Systems) received 21,575 submissions, ICLR 2025 (International Conference on Learning Representations) received 11,603 submissions, and ICML 2025 (International Conference on Machine Learning) received 12,107 (1-3). This volume has outpaced the supply of qualified human reviewers, and large language models (LLMs) are already filling the gap, largely undisclosed. An independent analysis of ICLR 2026 found that roughly 21% of the conference’s 75,800 peer reviews were fully AI (artificial intelligence)-generated, with over half showing some AI involvement, up from 15.8% at ICLR 2024 (4,5). This unregulated use carries documented risks: hallucinated citations have been found in accepted NeurIPS papers (6), and authors have embedded hidden prompt-injection instructions in manuscripts to manipulate AI reviewers into favorable assessments (7). We propose a triple-blind, multi-LLM pre-screening framework for peer review, developed for a health sciences journal, that formalizes and discloses AI involvement while preserving human peer reviewers as the final decision-making authority. The framework routes a submission through five stages — sanitization and anonymization, parallel AI pre-screening against a versioned rubric, an automated check gate, blinded human review, and editorial adjudication — with return-to-author loops at the check and editor stages. To address the confidentiality concerns that led NIH (National Institutes of Health) and NSF (National Science Foundation) to bar reviewers from submitting unpublished proposals to third-party generative AI (8,9), all three AI reviewers run on locally-hosted, open-weight LLMs, so manuscript content never leaves the journal’s own infrastructure. We situate the design against the closest existing precedent, Shen et al. (10), who benchmarked five opensource LLMs on single-pass quartile classification of 200 transplantation manuscripts and found accuracy insufficient (35% exact-match) for autonomous use — supporting, rather than undermining, our decision to retain mandatory human adjudication. This transparent, disclosed, human-supervised design offers a defensible alternative to today’s opaque, unregulated AI use in peer review, with the potential to reduce the substantial delay of traditional review (avg. 13 weeks / 91 days to first decision (11)) without displacing human judgment from the final decision.
Palavras-chave: peer review, artificial intelligence, large language models, triple-blind review, publication ethics, research integrity, academic publishing, editorial policy