[论文解读] It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
本文认为审前 RAI 数据集,特别是 COMPAS,存在偏见且在上下文上与公平性基准不匹配,现实世界的 CJ 结果取决于超越算法公平性的社会-技术因素。它倡导在使用 RAI 时采用跨学科标准和规范意识。
Risk assessment instrument (RAI) datasets, particularly ProPublica's COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance without accounting for the complexities of criminal justice (CJ) processes. However, we show that pretrial RAI datasets can contain numerous measurement biases and errors, and due to disparities in discretion and deployment, algorithmic fairness applied to RAI datasets is limited in making claims about real-world outcomes. These reasons make the datasets a poor fit for benchmarking under assumptions of ground truth and real-world impact. Furthermore, conventional practices of simply replicating previous data experiments may implicitly inherit or edify normative positions without explicitly interrogating value-laden assumptions. Without context of how interdisciplinary fields have engaged in CJ research and context of how RAIs operate upstream and downstream, algorithmic fairness practices are misaligned for meaningful contribution in the context of CJ, and would benefit from transparent engagement with normative considerations and values related to fairness, justice, and equality. These factors prompt questions about whether benchmarks for intrinsically socio-technical systems like the CJ system can exist in a beneficial and ethical way.
研究动机与目标
- 突出审前 RAI 数据集中的偏见与错误及其对基准公平性的影响。
- 解释为何仅仅依靠算法公平性不能在刑事司法语境中保障现实世界的公平。
- 在评估 RAI 时,鼓励来自犯罪学、心理学、法律与伦理学的跨学科标准。
- 为在 CJ 领域使用 RAI 数据集的研究提供规范性考量和实际指南。
提出的方法
- 对 Y(结果)、A(受保护属性)、X(协变量)及分布等方面的偏差与错误进行调查。
- 讨论刑事司法流程与自由裁量如何限制公平基准向现实世界结果的适用性。
- 将跨学科方法标准与标准机器学习基准实践进行比较。
- 在有情境性论证的前提下,提供使用 COMPAS/RAI 数据的建议和最佳实践。
- 批判性分析 AI 公平性实践与 CJ 研究规范之间的不匹配。
实验结果
研究问题
- RQ1审前 RAI 数据集中是否存在偏见和测量错误,从而削弱基准的有效性?
- RQ2刑事司法流程与人为裁量在何种程度上影响超越算法预测的现实世界公平?
- RQ3在研究 RAI 时,刑事司法领域的学科标准(犯罪学、心理学、法律)与机器学习基准规范有何不同?
- RQ4在公平性研究中使用 RAI 数据集应遵循哪些规范性考量?
- RQ5哪些最佳实践可以提升 COMPAS 与 RAI 数据在刑事司法研究中的有意义使用?
主要发现
- RAI 数据集(包括 COMPAS)在结果、受保护属性和协变量等方面存在测量偏差和错误,使基准评估变得复杂。
- 分布偏差、选择效应以及刑事司法流程中的下游裁量限制了将算法公平性结果转移到现实世界结果的力度。
- 简单重复先前的 RAI 实验可能在未明确审视公平性概念和 CJ 背景的情况下强化规范立场。
- 跨学科参与揭示,面向社会技术性 CJ 系统的基准需要情境、价值观和超越纯统计公平性的伦理考量。
- 当前的 ML 发表实践和以数据为中心的基准评估与 CJ 研究目标不一致,可能错误地呈现实世界影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。