[论文解读] Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments
本研究通过随机对照试验和观察性分析,调查了在NeurIPS 2022中评估同行评审质量的可靠性。研究发现存在显著偏差,尤其是在评审长度和作者结果方面,同时存在高水平的评审者间不一致、校准不当和主观性,表明人类对评审质量的评估不一致且存在缺陷,从而削弱了其在激励机制设计和干预效果评估中的应用价值。
Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed quality of reviews and measuring changes to review quality in experiments. We conduct a large scale study at the NeurIPS 2022 conference, a top-tier conference in machine learning, in which we invited (meta)-reviewers and authors to evaluate reviews given to submitted papers. First, we conduct a RCT to examine bias due to the length of reviews. We generate elongated versions of reviews by adding substantial amounts of non-informative content. Participants in the control group evaluate the original reviews, whereas participants in the experimental group evaluate the artificially lengthened versions. We find that lengthened reviews are scored (statistically significantly) higher quality than the original reviews. In analysis of observational data we find that authors are positively biased towards reviews recommending acceptance of their own papers, even after controlling for confounders of review length, quality, and different numbers of papers per author. We also measure disagreement rates between multiple evaluations of the same review of 28%-32%, which is comparable to that of paper reviewers at NeurIPS. Further, we assess the amount of miscalibration of evaluators of reviews using a linear model of quality scores and find that it is similar to estimates of miscalibration of paper reviewers at NeurIPS. Finally, we estimate the amount of variability in subjective opinions around how to map individual criteria to overall scores of review quality and find that it is roughly the same as that in the review of papers. Our results suggest that the various problems that exist in reviews of papers -- inconsistency, bias towards irrelevant factors, miscalibration, subjectivity -- also arise in reviewing of reviews.
研究动机与目标
- 评估在同行评审流程不同角色中,人类对评审质量评估的可靠性。
- 调查诸如评审长度或作者结果等偏差是否影响对评审质量的评估。
- 评估评审质量评估的一致性、校准程度和主观性。
- 探讨 flawed 的评审质量评估对同行评审中激励机制和实验干预的影响。
- 提供证据以判断当前评估评审质量的方法是否可作为研究同行评审的“金标准”。
提出的方法
- 在NeurIPS 2022中开展大规模随机对照试验,比较原始评审与人工延长但内容无信息量的版本的评估结果。
- 从评审人、副评审人和作者处收集对同一组评审的评估,以衡量评分者间不一致性和一致性。
- 使用Mann-Whitney U检验比较不同条件下的评估分数(例如,作者的“接受”与“拒绝”评审),同时控制长度和质量等混杂因素。
- 通过线性模型比较预测的评审质量分数与实际表现,来衡量校准不当程度。
- 通过估计从单个评分标准得分到整体评审质量得分的映射损失,量化主观性。
- 分析观察性数据,通过比较同一论文的“接受”与“拒绝”评审的评分,检测作者结果偏差。

实验结果
研究问题
- RQ1人为延长评审长度是否导致感知质量提高,表明存在无意义的冗长评审偏差?
- RQ2当评审建议接受其本人论文时,作者是否会给予更高评分,表明存在作者结果偏差?
- RQ3不同评审者对评审质量的评估在多大程度上保持一致?评审者间不一致的程度如何?
- RQ4评审者在评估评审质量时,其判断存在多大程度的校准不当?
- RQ5从单个评分标准得分映射到整体评审质量得分的过程中,主观性有多大?
主要发现
- 无意义的冗长评审偏差导致感知质量显著提高,等级双列相关系数为 τ = 0.64(p < 0.0001),延长后的评审在7分制量表上平均高出近0.5分。
- 作者对其本人论文推荐“接受”的评审表现出强烈的正面偏见,τ = 0.82(p < 0.0001),而对“拒绝”评审的评分平均低1.4分。
- 评审质量评估的评审者间不一致率在28%至32%之间,与NeurIPS中论文质量评估的不一致率相当。
- 评审者在评估评审质量时的校准不当程度与论文评审者相当,表明判断中存在系统性错误。
- 将评分标准得分映射到整体评审质量的主观性程度,与NeurIPS中论文评审评估所观察到的水平相当。
- 这些结果共同表明,论文评审中普遍存在的缺陷——偏差、不一致、校准不当和主观性——同样存在于对评审质量的评估中。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。