Skip to main content
QUICK REVIEW

[Paper Review] An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process

David Tran, Alexander V Valtchanov|arXiv (Cornell University)|Oct 11, 2020
scientometrics and bibliometrics research7 references19 citations
TL;DR

This study critically analyzes the ICLR 2017–2020 review process using public data, finding strong institutional bias and a significant gender gap in scores and acceptance rates, even after controlling for paper quality and topic. Despite moderate reproducibility (66% in 2020), reviewer scores poorly correlate with citation impact, and acceptance rates have declined over time, suggesting systemic inequities in peer review.

ABSTRACT

Mainstream machine learning conferences have seen a dramatic increase in the number of participants, along with a growing range of perspectives, in recent years. Members of the machine learning community are likely to overhear allegations ranging from randomness of acceptance decisions to institutional bias. In this work, we critically analyze the review process through a comprehensive study of papers submitted to ICLR between 2017 and 2020. We quantify reproducibility/randomness in review scores and acceptance decisions, and examine whether scores correlate with paper impact. Our findings suggest strong institutional bias in accept/reject decisions, even after controlling for paper quality. Furthermore, we find evidence for a gender gap, with female authors receiving lower scores, lower acceptance rates, and fewer citations per paper than their male counterparts. We conclude our work with recommendations for future conference organizers.

Motivation & Objective

  • To investigate reproducibility and randomness in the ICLR peer review process across 2017–2020.
  • To assess whether reviewer scores correlate with actual paper impact, measured by citations.
  • To examine whether the review process has deteriorated over time in terms of consensus, reproducibility, and correlation with impact.
  • To identify institutional bias in area chair decisions, particularly favoring prestigious institutions.
  • To investigate gender disparities in reviewer scores, acceptance rates, and citation impact.

Proposed method

  • Collected and curated data from OpenReview, arXiv, SemanticScholar, and institutional rankings (CS Rankings) for ICLR 2017–2020 papers.
  • Used Monte-Carlo simulations to quantify outcome reproducibility under varying reviewer counts and configurations.
  • Applied statistical models, including regression and non-parametric tests (e.g., Mann-Whitney U), to assess gender and institutional bias.
  • Categorized papers into 12 topics using hand-curated keywords to control for topic distribution in analyses.
  • Assigned gender labels to first and last authors using name-based and web-based heuristics, acknowledging limitations.
  • Measured citation impact using SemanticScholar data, with duplicate resolution via edit distance and manual curation.

Experimental results

Research questions

  • RQ1To what extent is the ICLR review process reproducible across different reviewer sets?
  • RQ2How well do reviewer scores predict actual citation impact of accepted papers?
  • RQ3Has the consistency and fairness of the review process declined over time (2017–2020)?
  • RQ4Is there evidence of institutional bias, where papers from prestigious institutions are more likely to be accepted regardless of reviewer scores?
  • RQ5Do gender disparities exist in reviewer scores, acceptance rates, and citation impact, even after controlling for topic and quality?

Key findings

  • The review process at ICLR showed 66% reproducibility in 2020, indicating moderate consistency, but this is inconsistent with the high perceived randomness by authors.
  • Reviewer scores showed only a weak correlation with actual citation impact, suggesting scores are poor predictors of long-term paper influence.
  • Papers from prestigious institutions had significantly higher acceptance rates even after controlling for reviewer scores, indicating strong institutional bias.
  • Female first authors received on average 0.16 points lower scores than male first authors (p = 0.085), and had a lower acceptance rate (22.1%) compared to male first authors (27.6%).
  • Even when controlling for topic distribution, women’s actual acceptance rate (22.1%) was lower than the expected rate (27.8%) if they had the same per-topic acceptance rates as men, indicating a systemic gap.
  • Women were underrepresented at ICLR: only 10.6% of first authors and 9.7% of last authors were female in 2020, despite 23.2% of CS PhD students being women in the US.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.