Skip to main content
QUICK REVIEW

[論文レビュー] Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments

Alexander Goldberg, Ivan Stelmakh|arXiv (Cornell University)|Nov 16, 2023
Expert finding and Q&A systems被引用数 7
ひとこと要約

本研究は、NeurIPS 2022における査読の質を評価する際の信頼性を、無作為化比較試験と観察的分析を用いて調査した。査読の長さや著者による結果に顕著なバイアスが存在し、査読者間の合意度が低く、校正が不適切で主観的であることが判明した。これは、人間による査読の質の評価が一貫性がなく、欠陥があることを示しており、インcentive設計や介入評価への応用を損なうものである。

ABSTRACT

Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed quality of reviews and measuring changes to review quality in experiments. We conduct a large scale study at the NeurIPS 2022 conference, a top-tier conference in machine learning, in which we invited (meta)-reviewers and authors to evaluate reviews given to submitted papers. First, we conduct a RCT to examine bias due to the length of reviews. We generate elongated versions of reviews by adding substantial amounts of non-informative content. Participants in the control group evaluate the original reviews, whereas participants in the experimental group evaluate the artificially lengthened versions. We find that lengthened reviews are scored (statistically significantly) higher quality than the original reviews. In analysis of observational data we find that authors are positively biased towards reviews recommending acceptance of their own papers, even after controlling for confounders of review length, quality, and different numbers of papers per author. We also measure disagreement rates between multiple evaluations of the same review of 28%-32%, which is comparable to that of paper reviewers at NeurIPS. Further, we assess the amount of miscalibration of evaluators of reviews using a linear model of quality scores and find that it is similar to estimates of miscalibration of paper reviewers at NeurIPS. Finally, we estimate the amount of variability in subjective opinions around how to map individual criteria to overall scores of review quality and find that it is roughly the same as that in the review of papers. Our results suggest that the various problems that exist in reviews of papers -- inconsistency, bias towards irrelevant factors, miscalibration, subjectivity -- also arise in reviewing of reviews.

研究の動機と目的

  • 査読プロセスにおける異なる役割における人間による査読の質の評価の信頼性を評価すること。
  • 査読の長さや著者による結果といったバイアスが、査読の質の評価に影響を与えるかどうかを調査すること。
  • 査読の質の評価の一貫性、校正、主観性を評価すること。
  • 査読の質の評価に欠陥があることが、査読におけるインcentiveメカニズムや実験的介入に与える影響を検討すること。
  • 現在の査読の質を評価する方法が、査読に関する研究における「ゴールドスタンダード」として信頼できるかどうかの証拠を提供すること。

提案手法

  • NeurIPS 2022において大規模な無作為化比較試験を実施し、オリジナルの査読と非情報的コンテンツを含む人工的に長くされたバージョンの査読を比較した。
  • 同じ査読セットについて、査読者、メタ査読者、著者が評価を提供した。これにより、評価者間の不一致と一貫性を評価した。
  • 長さや質といった交絡要因を調整した上で、Mann-Whitney U検定を用いて、条件別(例:著者が「受諾」を推奨した査読 vs. 「却下」を推奨した査読)の評価スコアを比較した。
  • 線形モデルを用いて予測された査読の質スコアと実際のパフォーマンスを比較することで、校正の不適切さを測定した。
  • 個々の基準スコアから全体の査読の質スコアへのマッピングの学習モデルの損失を推定することで、主観性を定量化した。
  • 観察的データを分析し、同じ論文の「受諾」査読と「却下」査読の評価を比較することで、著者による結果バイアスを検出した。
(a) Overall review quality score.
(a) Overall review quality score.

実験結果

リサーチクエスチョン

  • RQ1人工的に査読の長さを延ばすと、評価された質が上昇するか。これは、意味のない長さの延長によるバイアスを示唆するか。
  • RQ2著者が自身の論文を受諾するよう勧める査読に対して、著者がより好意的に評価するか。これは、著者による結果バイアスを示唆するか。
  • RQ3異なる評価者間で査読の質の評価はどれほど一貫しているのか。評価者間の不一致の程度はどの程度か。
  • RQ4評価者が査読の質をどの程度校正が不適切に評価しているか。
  • RQ5個々の基準スコアから全体の査読の質スコアへのマッピングに、どの程度の主観性が存在するか。

主な発見

  • 意味のない長さの延長によるバイアスにより、評価された質が有意に上昇し、ランク・バイセラル相関係数 τ = 0.64(p < 0.0001)を示した。これにより、長くされた査読は7段階スケールで平均0.5ポイント高いスコアを獲得した。
  • 著者は自身の論文を受諾するよう勧める査読に対して強い肯定バイアスを示し、τ = 0.82(p < 0.0001)であった。一方、「却下」査読は平均1.4ポイント低い評価を受けた。
  • 査読の質に関する評価者間の不一致率は28%〜32%であり、NeurIPSにおける論文の質の評価で観察された不一致率と同等であった。
  • 評価者が査読の質を評価する際の校正の不適切さは、論文査読者と同程度であり、判断に系統的な誤りが存在することを示した。
  • 個々の基準スコアから全体の査読の質スコアへのマッピングにおける主観性は、NeurIPSにおける論文査読評価で観察されたものと定量的に同等であった。
  • これらの結果は、論文査読に見られるのと同じ欠陥—バイアス、一貫性の欠如、校正の不適切さ、主観性—が、査読の質の評価に対しても同様に存在することを総合的に示している。
(b) Criteria scores.
(b) Criteria scores.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。