Skip to main content
QUICK REVIEW

[论文解读] ProPublica's COMPAS Data Revisited

Matias Barenstein|arXiv (Cornell University)|Jun 11, 2019
Ethics and Social Impacts of AI参考文献 6被引用 5
一句话总结

本文揭示了普罗布里卡(ProPublica)广泛使用的COMPAS再犯数据集中存在一个关键的数据处理错误:普罗布里卡未对再犯者应用两年筛查日期的截止规则,导致两年再犯率被高估了24.3%(从36.2%上升至45.1%)。作者通过统一应用该截止规则对数据进行更正,表明关键公平性指标如PPV和NPV受到显著影响,而FPR、FNR和准确率则基本保持不变。

ABSTRACT

I examine the COMPAS recidivism risk score and criminal history data collected by ProPublica in 2016 that fueled intense debate and research in the nascent field of 'algorithmic fairness'. ProPublica's COMPAS data is used in an increasing number of studies to test various definitions of algorithmic fairness. This paper takes a closer look at the actual datasets put together by ProPublica. In particular, the sub-datasets built to study the likelihood of recidivism within two years of a defendant's original COMPAS survey screening date. I take a new yet simple approach to visualize these data, by analyzing the distribution of defendants across COMPAS screening dates. I find that ProPublica made an important data processing error when it created these datasets, failing to implement a two-year sample cutoff rule for recidivists in such datasets (whereas it implemented a two-year sample cutoff rule for non-recidivists). When I implement a simple two-year COMPAS screen date cutoff rule for recidivists, I estimate that in the two-year general recidivism dataset ProPublica kept over 40% more recidivists than it should have. This fundamental problem in dataset construction affects some statistics more than others. It obviously has a substantial impact on the recidivism rate; artificially inflating it. For the two-year general recidivism dataset created by ProPublica, the two-year recidivism rate is 45.1%, whereas, with the simple COMPAS screen date cutoff correction I implement, it is 36.2%. Thus, the two-year recidivism rate in ProPublica's dataset is inflated by over 24%. This also affects the positive and negative predictive values. On the other hand, this data processing error has little impact on some of the other key statistical measures, which are less susceptible to changes in the relative share of recidivists, such as the false positive and false negative rates, and the overall accuracy.

研究动机与目标

  • 调查普罗布里卡广泛使用的COMPAS再犯数据集的完整性,该数据集支撑了大量算法公平性研究。
  • 识别并纠正构建两年再犯数据集过程中存在的关键数据处理错误。
  • 评估该错误对关键公平性指标(如再犯率、PPV、NPV、FPR/FNR)的影响。
  • 强调在算法公平性研究的基准数据集中,数据质量和处理透明度的重要性。
  • 提供经更正的数据集,并警示在未对数据构建过程进行严格审查的情况下使用数据集所存在的风险。

提出的方法

  • 作者分析被告在COMPAS筛查日期上的分布,以检测数据处理中的不一致之处。
  • 对再犯者和非再犯者统一应用两年筛查日期截止规则,纠正普罗布里卡原始处理中存在的一致性偏差。
  • 使用更正后的数据集重新计算关键统计指标,包括再犯率、PPV、NPV、FPR、FNR和准确率。
  • 通过对比原始普罗布里卡统计数据与更正后数据集的统计结果,量化该更正的影响。
  • 分析聚焦于两年期一般再犯数据集,这是公平性研究中最常使用的子集。
  • 作者通过可视化和统计分析表明,原始数据集中再犯者比例异常偏高,原因是未对再犯者应用截止规则。

实验结果

研究问题

  • RQ1普罗布里卡在其两年再犯数据集中,是否对再犯者和非再犯者统一应用了两年筛查日期截止规则?
  • RQ2缺失的截止规则对两年再犯率的量化影响是什么?
  • RQ3该数据处理错误如何影响PPV和NPV等关键公平性指标?
  • RQ4该数据错误对假正类率(FPR)、假负类率(FNR)和准确率的影响程度如何?
  • RQ5该数据处理错误对使用该数据集进行算法公平性研究有效性的更广泛影响是什么?

主要发现

  • 普罗布里卡未对再犯者应用两年筛查日期截止规则,而对非再犯者应用了该规则,导致再犯者在两年数据集中系统性地被过度代表。
  • 普罗布里卡原始数据集中两年再犯率被高估了8.8个百分点,从36.2%上升至45.1%。
  • 这相当于由于数据处理错误,两年再犯率相对上升了24.3%。
  • 由于PPV和NPV依赖于再犯者在总体中的相对比例,该错误对这两个指标产生了显著影响。
  • 假正类率(FPR)、假负类率(FNR)和整体准确率基本不受影响,因为它们对再犯者比例的变化不敏感。
  • 作者估计,由于未对再犯者应用截止规则,普罗布里卡多保留了44.3%的两年再犯者。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。