[论文解读] AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews
本文提出了三种基于 LLM 的审稿系统(OpenReviewer、Papers with Reviews、Reviewer Arena)以及四种评估方法,用以评估与人类偏好的一致性、偏见及大规模学术评审的局限性。
Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena of human preferences by pairwise comparisons. Gathering human preference may be time-consuming; therefore, we also use an LLM to automatically evaluate reviews to increase sample efficiency while reducing bias. In addition to evaluating human and LLM preferences among LLM reviews, we fine-tune an LLM to predict human preferences, predicting which reviews humans will prefer in a head-to-head battle between LLMs. We artificially introduce errors into papers and analyze the LLM's responses to identify limitations, use adaptive review questions, meta prompting, role-playing, integrate visual and textual analysis, use venue-specific reviewing materials, and predict human preferences, improving upon the limitations of the traditional review processes. We make the reviews of publicly available arXiv and open-access Nature journal papers available online, along with a free service which helps authors review and revise their research papers and improve their quality. This work develops proof-of-concept LLM reviewing systems that quickly deliver consistent, high-quality reviews and evaluate their quality. We mitigate the risks of misuse, inflated review scores, overconfident ratings, and skewed score distributions by augmenting the LLM with multiple documents, including the review form, reviewer guide, code of ethics and conduct, area chair guidelines, and previous year statistics, by finding which errors and shortcomings of the paper may be detected by automated reviews, and evaluating pairwise reviewer preferences. This work identifies and addresses the limitations of using LLMs as reviewers and evaluators and enhances the quality of the reviewing process.
研究动机与目标
- 在大规模基础模型辅助的审稿中体现动机,降低偏见,同时保持质量控制。
- 开发并部署三种审稿系统,以生成、收集并评估 arXiv 与开放获取 Nature 论文的评稿。
- 通过人类偏好、自动化的 LLM 评估和偏好预测,评估 LLM 审稿与人类审稿之间的一致性。
- 识别基于 LLM 的审稿的局限性和潜在风险,并提出缓解策略。
提出的方法
- 三种审稿系统:OpenReviewer(LLM 辅助评审)、Papers with Reviews(大规模评审收集与打分)、Reviewer Arena(评审之间的成对比较)。
- 四种评估方法:匿名人工评估、自动化 LLM 评估、自动化的 LLM 对人类偏好的预测,以及通过有意论文修改自动发现 LLM 审稿的局限性。
- LLMs 进行角色扮演,模拟人类编辑流程(作者、评审、领域主席、项目主席)。
- 将多份文献(评审表、指南、伦理规范、统计数据)作为上下文,以校准 LLM 审稿并使其与场刊规范保持一致。

实验结果
研究问题
- RQ1LLM 生成的评审在盲评中的人类偏好以及基于 GPT-4 的比较中,能否与人类偏好保持一致?
- RQ2在固定、可适应与生成的评审提示下,LLMs 作为学术评审者的优点与局限性是什么?
- RQ3成对偏好数据、BT 建模与自动评估方法如何量化评审质量与排名?
- RQ4LLM 基于审稿所呈现的偏见与错误有哪些,如何通过提示、上下文和后处理进行缓解?
- RQ5场刊特定指南和补充材料如何影响自动评审的质量与可信度?
主要发现
- LLM 审稿在盲评和基于 GPT-4 的比较中与人类评审保持了相当程度的一致性,在某些情境下,某些模型甚至优于人类。
- GPT-4 Turbo(April 9, 2024)在五位评审的人工偏好测试中获得最高排名;人类排名第二,其次是其他 LLM。
- Bradley-Terry 建模给出评审者的强度排序;GPT-4 Turbo 位列第一,人类次之,然后是 Command R+,Claude 3 Opus 和 Gemini Pro 落后。
- 基于 PPI 的自动评估方法可以减少对人类数据的依赖并提高偏好预测的效率。
- 通过引入论文错误来自动发现局限性,有助于映射 LLM 对特定内容类型与不足之处的敏感性。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。