Using Artificial Intelligence in Scholarly Review: Reconciling the Limits of Human Judgment with the Limits of the Machine
- Debra E. Henninger-Borckardt, PhD — Medical University of South Carolina
Abstract
Peer review remains the primary quality-control mechanism of scholarly publishing, yet the evidence based on human reviewer performance shows chronically low inter-rater reliability, documented bias, and a labor system straining under rising submission volumes. The emergence of large language models (LLMs) capable of generating manuscript feedback has prompted journals, including those in health professions education, to ask whether artificial intelligence (AI) can serve as a reviewer, an aid to reviewers, or neither. This article synthesizes the empirical literature on human reviewer reliability and bias, catalogs the documented strengths and vulnerabilities of LLM-generated review, and evaluates the case for hybrid human-AI review systems. Across multiple large-scale studies, AI-generated feedback overlaps with human reviewer commentary at rates statistically comparable to the overlap between two independent human reviewers, and AI systems can process volume, apply criteria consistently, and surface omissions humans overlook. At the same time, AI reviewers hallucinate citations, are demonstrably manipulable through prompt injection embedded in manuscripts, and lack the contextual and ethical judgment that anchors legitimate peer review. The data available at present support neither wholesale automation nor blanket prohibition, but rather a structured hybrid model in which AI performs triage, consistency-checking, and completeness screening under human oversight, while accountability for the final editorial judgment remains with human reviewers and editors.
How to cite
Open access under CC BY 4.0. © 2026 the author(s).