[Paper Review] AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews
The paper presents three LLM-based review systems (OpenReviewer, Papers with Reviews, Reviewer Arena) and four evaluation methods to assess alignment with human preferences, bias, and limitations in scalable academic reviewing.
Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena of human preferences by pairwise comparisons. Gathering human preference may be time-consuming; therefore, we also use an LLM to automatically evaluate reviews to increase sample efficiency while reducing bias. In addition to evaluating human and LLM preferences among LLM reviews, we fine-tune an LLM to predict human preferences, predicting which reviews humans will prefer in a head-to-head battle between LLMs. We artificially introduce errors into papers and analyze the LLM's responses to identify limitations, use adaptive review questions, meta prompting, role-playing, integrate visual and textual analysis, use venue-specific reviewing materials, and predict human preferences, improving upon the limitations of the traditional review processes. We make the reviews of publicly available arXiv and open-access Nature journal papers available online, along with a free service which helps authors review and revise their research papers and improve their quality. This work develops proof-of-concept LLM reviewing systems that quickly deliver consistent, high-quality reviews and evaluate their quality. We mitigate the risks of misuse, inflated review scores, overconfident ratings, and skewed score distributions by augmenting the LLM with multiple documents, including the review form, reviewer guide, code of ethics and conduct, area chair guidelines, and previous year statistics, by finding which errors and shortcomings of the paper may be detected by automated reviews, and evaluating pairwise reviewer preferences. This work identifies and addresses the limitations of using LLMs as reviewers and evaluators and enhances the quality of the reviewing process.
Motivation & Objective
- Motivate the need for foundation-model-assisted reviewing at scale and reduce bias while maintaining quality controls.
- Develop and deploy three review systems to generate, collect, and evaluate reviews for arXiv and open-access Nature papers.
- Evaluate alignment between LLM reviews and human reviews using human preferences, automatic LLM evaluation, and preference prediction.
- Identify limitations and potential risks of LLM-based reviewing and propose mitigation strategies.
Proposed method
- Three review systems: OpenReviewer (LLM-assisted reviews), Papers with Reviews (large-scale review collection and scoring), and Reviewer Arena (pairwise comparison of reviews).
- Four evaluation methods: anonymous human evaluation, automatic LLM evaluation, automatic LLM prediction of human preferences, and automatic discovery of LLM review limitations via deliberate paper modifications.
- Role-playing by LLMs to simulate human editorial processes (authors, reviewers, area chairs, program chairs).
- Use of multiple documents (review form, guidelines, ethics codes, statistics) as context to calibrate LLM reviews and align with venue norms.

Experimental results
Research questions
- RQ1Can LLM-generated reviews align with human preferences in blind evaluations and in GPT-4-based comparisons?
- RQ2What are the strengths and limitations of LLMs as academic reviewers across fixed, adaptive, and generated review prompts?
- RQ3How can pairwise preference data, BT modeling, and autoevaluation approaches quantify reviewer quality and ranking?
- RQ4What biases and errors do LLM-based reviews exhibit, and how can they be mitigated through prompting, context, and post-processing?
- RQ5How do venue-specific guidelines and supplementary materials influence the quality and trustworthiness of automated reviews?
Key findings
- LLM reviews align reasonably with human reviews in blind evaluations and GPT-4-based comparisons, with certain models outperforming humans in some settings.
- GPT-4 Turbo (April 9, 2024) achieves the top ranking in human-preference tests among five reviewers; humans rank second, followed by other LLMs.
- Bradley-Terry modeling yields a ranked strength order of reviewers; GPT-4 Turbo leads, followed by Human, then Command R+, with Claude 3 Opus and Gemini Pro lagging.
- Automatic evaluation using PPI-based methods can reduce reliance on human data and improve efficiency in preference prediction.
- Automatic discovery of limitations by introducing paper errors helps map LLM sensitivity to specific types of content and shortcomings.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.