Skip to main content
QUICK REVIEW

[论文解读] On Testing for Biases in Peer Review

Ivan Stelmakh, Nihar B. Shah|arXiv (Cornell University)|Dec 31, 2019
Expert finding and Q&A systems被引用 15
一句话总结

本文指出了Tomkins等人在大规模同行评审实验中所采用的统计检验方法中的关键缺陷,表明其检验在存在评审员噪声、模型不匹配和校准问题等现实条件下的情况下,无法控制第一类错误(假阳性)。本文提出了一种新颖且稳健的假设检验框架——基于分歧与计数检验——在最小假设条件下,即使存在细微偏差或测量误差,也能保证可控的误报率和非平凡的检验效能。

ABSTRACT

We consider the issue of biases in scholarly research, specifically, in peer review. There is a long standing debate on whether exposing author identities to reviewers induces biases against certain groups, and our focus is on designing tests to detect the presence of such biases. Our starting point is a remarkable recent work by Tomkins, Zhang and Heavlin which conducted a controlled, large-scale experiment to investigate existence of biases in the peer reviewing of the WSDM conference. We present two sets of results in this paper. The first set of results is negative, and pertains to the statistical tests and the experimental setup used in the work of Tomkins et al. We show that the test employed therein does not guarantee control over false alarm probability and under correlations between relevant variables coupled with any of the following conditions, with high probability, can declare a presence of bias when it is in fact absent: (a) measurement error, (b) model mismatch, (c) reviewer calibration. Moreover, we show that the setup of their experiment may itself inflate false alarm probability if (d) bidding is performed in non-blind manner or (e) popular reviewer assignment procedure is employed. Our second set of results is positive and is built around a novel approach to testing for biases that we propose. We present a general framework for testing for biases in (single vs. double blind) peer review. We then design hypothesis tests that under minimal assumptions guarantee control over false alarm probability and non-trivial power even under conditions (a)--(c) as well as propose an alternative experimental setup which mitigates issues (d) and (e). Finally, we show that no statistical test can improve over the non-parametric tests we consider in terms of the assumptions required to control for the false alarm probability.

研究动机与目标

  • 批判性评估Tomkins等人在大规模WSDM同行评审实验中所用假设检验的统计有效性。
  • 证明原始检验在存在评审员噪声和模型不匹配等常见现实条件时,无法控制第一类错误(假阳性率)。
  • 开发一种用于检测单盲与双盲同行评审中偏见的新统计框架,确保可控的误报率和非平凡的检测效能。
  • 提出一种改进的实验设计,以减少因非盲投标和常见评审员分配程序导致的假阳性膨胀。
  • 在控制第一类错误所需假设的条件下,建立所提非参数检验的理论最优性。

提出的方法

  • 提出一种基于分歧的检验方法,通过比较单盲与双盲条件下评审员评分的差异,以检测系统性偏见。
  • 引入一种基于计数的检验方法,通过聚合不同条件下评审员的决策结果,检测在原假设下行为的偏离。
  • 采用广义逻辑斯蒂模型,形式化在原假设(无偏见)和备择假设(存在偏见)下评审员决策的概率。
  • 使用随机分配程序,确保在原假设下所有可能的评审员-论文分配组合具有相等的概率,从而支持有效的统计推断。
  • 应用全概率定律,将条件结果推广至所有可能配置下的无条件第一类错误控制。
  • 证明所提出的非参数检验在控制第一类错误所需假设方面具有理论最优性:任何其他检验都无法在更弱假设下实现更好的控制。

实验结果

研究问题

  • RQ1Tomkins等人在WSDM实验中所用的统计检验在存在评审员噪声和模型不匹配等现实条件时,能否可靠地控制第一类错误?
  • RQ2能否设计一种新的假设检验框架,使其在最小假设条件下,保证第一类错误控制和非平凡的检验效能?
  • RQ3非盲投标和标准评审员分配程序等因素如何影响偏见检测中的假阳性率?
  • RQ4在控制第一类错误所需假设的条件下,任何统计检验在偏见检测中的性能理论极限是什么?
  • RQ5当存在测量误差或评审员校准偏差导致评分失真时,所提检验能否检测到细微偏见?

主要发现

  • Tomkins等人研究中所用的检验在合理条件下(如评审员噪声和模型不匹配)无法控制第一类错误,假阳性率高达0.5或更高。
  • 所提出的分歧检验与计数检验在最小假设下均能保证第一类错误控制,即使存在测量误差、模型不匹配或评审员校准问题。
  • 新的实验设置可有效缓解因非盲投标和常见评审员分配程序导致的误报膨胀。
  • 所提非参数检验具有理论最优性:没有任何其他检验能在更弱假设下实现更好的第一类错误控制。
  • 模拟结果表明,所提检验在评审员评分相关性和估计噪声条件下仍保持高检测效能,优于原始方法。
  • 该框架对强参数假设的违反具有鲁棒性,适用于存在评审员行为异质性的现实同行评审场景。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。