[Paper Review] It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
The paper argues that pretrial RAI datasets, notably COMPAS, are biased and contextually misaligned for benchmarking fairness, and that real-world CJ outcomes depend on socio-technical factors beyond algorithmic fairness. It advocates interdisciplinary standards and normative awareness when using RAIs.
Risk assessment instrument (RAI) datasets, particularly ProPublica's COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance without accounting for the complexities of criminal justice (CJ) processes. However, we show that pretrial RAI datasets can contain numerous measurement biases and errors, and due to disparities in discretion and deployment, algorithmic fairness applied to RAI datasets is limited in making claims about real-world outcomes. These reasons make the datasets a poor fit for benchmarking under assumptions of ground truth and real-world impact. Furthermore, conventional practices of simply replicating previous data experiments may implicitly inherit or edify normative positions without explicitly interrogating value-laden assumptions. Without context of how interdisciplinary fields have engaged in CJ research and context of how RAIs operate upstream and downstream, algorithmic fairness practices are misaligned for meaningful contribution in the context of CJ, and would benefit from transparent engagement with normative considerations and values related to fairness, justice, and equality. These factors prompt questions about whether benchmarks for intrinsically socio-technical systems like the CJ system can exist in a beneficial and ethical way.
Motivation & Objective
- Highlight biases and errors in pretrial RAI datasets and their impact on benchmarking fairness.
- Explain why algorithmic fairness alone cannot guarantee real-world fairness in criminal justice contexts.
- Encourage interdisciplinary standards from criminology, psychology, law, and ethics when evaluating RAIs.
- Provide normative considerations and practical guidelines for research using RAI datasets in CJ.
Proposed method
- Survey biases and errors across Y (outcomes), A (protected attributes), X (covariates), and distribution.
- Discuss how CJ processes and discretion limit the applicability of fairness benchmarks to real-world outcomes.
- Compare interdisciplinary methodological standards with standard ML benchmarking practices.
- Offer recommendations and best practices for using COMPAS/RAI data with contextual justification.
- Critically analyze the mismatch between AI fairness practices and CJ research norms.
Experimental results
Research questions
- RQ1Do pretrial RAI datasets contain biases and measurement errors that undermine benchmarking validity?
- RQ2In what ways do CJ processes and human discretion affect real-world fairness beyond algorithmic predictions?
- RQ3How do disciplinary standards in CJ (criminology, psychology, law) differ from ML benchmarking norms when studying RAIs?
- RQ4What normative considerations should guide the use of RAI datasets in fairness research?
- RQ5What best practices can improve the meaningful use of COMPAS and RAI data in CJ research?
Key findings
- RAI datasets, including COMPAS, contain measurement biases and errors across outcomes, protected attributes, and covariates that complicate benchmarking.
- Distributional biases, selection effects, and downstream discretion in CJ processes limit the transfer of algorithmic fairness results to real-world outcomes.
- Simply replicating previous RAI experiments can reinforce normative positions without explicit interrogation of fairness concepts and CJ context.
- Interdisciplinary engagement reveals that benchmarks for socio-technical CJ systems require context, values, and ethical considerations beyond pure statistical fairness.
- Current ML publication practices and data-centric benchmarking misalign with CJ research goals and can misrepresent real-world impact.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.