[Paper Review] Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
The paper shows that large language models used for scholarly peer review are vulnerable to explicit and implicit manipulation, have inherent flaws and biases, and thus are not ready for widespread adoption.
Scholarly peer review is a cornerstone of scientific advancement, but the system is under strain due to increasing manuscript submissions and the labor-intensive nature of the process. Recent advancements in large language models (LLMs) have led to their integration into peer review, with promising results such as substantial overlaps between LLM- and human-generated reviews. However, the unchecked adoption of LLMs poses significant risks to the integrity of the peer review system. In this study, we comprehensively analyze the vulnerabilities of LLM-generated reviews by focusing on manipulation and inherent flaws. Our experiments show that injecting covert deliberate content into manuscripts allows authors to explicitly manipulate LLM reviews, leading to inflated ratings and reduced alignment with human reviews. In a simulation, we find that manipulating 5% of the reviews could potentially cause 12% of the papers to lose their position in the top 30% rankings. Implicit manipulation, where authors strategically highlight minor limitations in their papers, further demonstrates LLMs' susceptibility compared to human reviewers, with a 4.5 times higher consistency with disclosed limitations. Additionally, LLMs exhibit inherent flaws, such as potentially assigning higher ratings to incomplete papers compared to full papers and favoring well-known authors in single-blind review process. These findings highlight the risks of over-reliance on LLMs in peer review, underscoring that we are not yet ready for widespread adoption and emphasizing the need for robust safeguards.
Motivation & Objective
- Motivate the strain on traditional peer review due to rising submissions and labor demands.
- Assess whether LLMs can reliably review scholarly manuscripts.
- Identify manipulation vectors (explicit and implicit) that can sway LLM reviews.
- Investigate inherent flaws and biases in LLM-based reviews such as hallucination, length bias, and author bias.
Proposed method
- Replicate three established LLM-based reviewing pipelines linked to human alignment with human reviews.
- Develop explicit manipulation via invisible white-text injections in manuscripts to steer LLM reviews toward acceptance.
- Examine implicit manipulation by analyzing authors highlighting limitations and its effect on LLM vs human reviews.
- Assess inherent flaws including hallucinations with incomplete content, length bias, and authorship bias across multiple LLMs.
- Quantify effects using consistency metrics between LLM and human reviews and a rating-to-paper model to simulate decision impact.
Experimental results
Research questions
- RQ1Can LLM-based reviews be manipulated to diverge from human judgments through covert input in manuscripts?
- RQ2Do authors’ disclosed limitations bias LLM reviews more than human reviews?
- RQ3What inherent flaws or biases do LLMs exhibit (e.g., hallucination, length, authorship) in the peer-review setting?
- RQ4How could manipulated LLM reviews affect paper rankings and acceptance decisions?
Key findings
- Explicit manipulation can drastically reduce LLM–human review consistency (e.g., from 53.29 to 15.91).
- Manipulated content injected into manuscripts can cause LLM reviews to align with the injected content at high rates ( Injection–LLM-Matched / Injection rises to 92.49%).
- Five percent of manipulated reviews could cause about 12% of papers to drop out of the top 30% rankings.
- LLMs are 4.5× more consistent with authors’ proclaimed limitations than humans, indicating vulnerability to implicit manipulation.
- LLMs can hallucinate with incomplete inputs (e.g., empty papers) and may rate incomplete papers similarly to full ones, revealing unreliability in using LLMs for reviews.
- In single-blind settings, LLMs show bias toward well-known authors or affiliations, suggesting fairness concerns.
- LLM performance in consistency with human reviews correlates with overall model capabilities (e.g., GPT-4o-0806 strongest among tested models).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.