Skip to main content
QUICK REVIEW

[论文解读] Arbitrariness of peer review: A Bayesian analysis of the NIPS experiment

Olivier François|arXiv (Cornell University)|Jul 23, 2015
Explainable Artificial Intelligence (XAI)参考文献 13被引用 12
一句话总结

本文应用贝叶斯统计建模分析NIPS同行评审实验,引入一个隐藏参数以表示符合基本质量标准的投稿比例。研究发现56%的投稿符合质量标准(95%可信区间:0.34–0.83),若将接受率提高至该水平,可降低同行评审中的随意性。

ABSTRACT

The principle of peer review is central to the evaluation of research, by ensuring that only high-quality items are funded or published. But peer review has also received criticism, as the selection of reviewers may introduce biases in the system. In 2014, the organizers of the ``Neural Information Processing Systems q q{} conference conducted an experiment in which $10\%$ of submitted manuscripts (166 items) went through the review process twice. Arbitrariness was measured as the conditional probability for an accepted submission to get rejected if examined by the second committee. This number was equal to $60\%$, for a total acceptance rate equal to $22.5\%$. Here we present a Bayesian analysis of those two numbers, by introducing a hidden parameter which measures the probability that a submission meets basic quality criteria. The standard quality criteria usually include novelty, clarity, reproducibility, correctness and no form of misconduct, and are met by a large proportions of submitted items. The Bayesian estimate for the hidden parameter was equal to $56\%$ ($95\%$CI: $ I = (0.34, 0.83)$), and had a clear interpretation. The result suggested the total acceptance rate should be increased in order to decrease arbitrariness estimates in future review processes.

研究动机与目标

  • 使用贝叶斯推断分析NIPS实验数据,量化同行评审中的随意性。
  • 估计符合基本质量标准(如新颖性、正确性、可复现性)的投稿的隐藏比例。
  • 评估两阶段评审模型——先检查质量,再随机接受——是否比单阶段模型更能解释观测数据。
  • 评估提高总体接受率是否可降低未来同行评审过程中的随意性。
  • 比较不同关于高质量投稿分布假设下的模型拟合度。

提出的方法

  • 提出一种拒绝或抛硬币(Reject or Flip a Coin, RFC)模型,包含隐藏参数x(符合基本质量标准的投稿比例)和y(在满足质量标准条件下的接受概率)。
  • 使用非信息先验,通过贝叶斯定理推导后验分布,将随意性建模为a = 1 - π/x。
  • 引入扩展的RAFC模型,允许高质量投稿占比例α,通过近似贝叶斯计算(ABC)估计α。
  • 使用100,000次模拟进行ABC,从后验分布中抽样并计算不同α值下的模型概率。
  • 利用后验分布计算x和y的可信区间,报告95%可信区间。
  • 通过后验模型概率和贝叶斯因子比较模型拟合度,评估RAFC模型是否优于RFC模型。

实验结果

研究问题

  • RQ1有多少比例的投稿符合基本质量标准?该估计的不确定性如何?
  • RQ2观测到的60%随意性与投稿的潜在质量水平有何关系?
  • RQ3两阶段评审模型——先筛选质量,再随机接受——是否比单阶段模型更能解释NIPS数据?
  • RQ4α(高质量投稿比例)的最佳拟合值是多少?它如何影响模型性能?
  • RQ5若将总体接受率提高到与估计的质量阈值一致,是否可降低同行评审中的随意性?

主要发现

  • 符合基本质量标准的投稿比例的贝叶斯估计为56%,95%可信区间为(0.34, 0.83)。
  • 观测到的60%随意性与仅56%的投稿满足最低质量标准的模型一致。
  • 在测试的α值中,α = 5%的RAFC模型拟合度最佳,后验模型概率为26%。
  • 与RFC模型的成对比较得到贝叶斯因子1.31,表明对RAFC模型仅有轻微证据支持。
  • 若将总体接受率提高至56%,理论上可使随意性降至接近零,前提是质量标准得到满足。
  • 结果表明,通过将质量筛选与最终选择分离,可简化同行评审流程,减轻评审者负担并降低随意性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。