Skip to main content
QUICK REVIEW

[Paper Review] Fair prediction with disparate impact: A study of bias in recidivism prediction instruments

Alexandra Chouldechova|arXiv (Cornell University)|Oct 24, 2016
Regulation and Compliance StudiesBusiness, Management and Accounting18 citations
TL;DR

This paper demonstrates that recidivism prediction instruments (RPIs) can produce disparate impact—even when they are free from predictive bias—due to differing recidivism rates across racial groups. It shows that well-calibrated RPIs like COMPAS may still result in higher false positive rates for Black defendants, leading to unequal incarceration outcomes under binary penalty policies.

ABSTRACT

Recidivism prediction instruments provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the country, their use is attracting tremendous controversy. Much of the controversy concerns potential discriminatory bias in the risk assessments that are produced. This paper discusses a fairness criterion originating in the field of educational and psychological testing that has recently been applied to assess the fairness of recidivism prediction instruments. We demonstrate how adherence to the criterion may lead to considerable disparate impact when recidivism prevalence differs across groups.

Motivation & Objective

  • To investigate how fairness criteria in psychometric testing interact with real-world outcomes in recidivism prediction.
  • To examine the conditions under which well-calibrated RPIs lead to disparate impact across racial groups.
  • To clarify the statistical link between test fairness, error rates, and policy outcomes in risk assessment.
  • To evaluate whether balancing error rates across groups is sufficient to prevent unfair outcomes at finer subgroups.
  • To argue against abandoning data-driven risk assessment due to headline-grabbing bias reports, advocating for context-aware fairness evaluation instead.

Proposed method

  • Uses the Broward County COMPAS dataset with race (Black/White), decile scores, and 2-year recidivism outcomes.
  • Defines test fairness via well-calibration: P(Y=1|S=s, R=r) is equal across racial groups for all score values s.
  • Introduces a coarsened risk score S_c to classify individuals as high-risk (HR) or low-risk (LR) based on a threshold s_HR.
  • Analyzes error rates (FPR, FNR) and positive predictive value (PPV) across racial groups under the fairness constraint.
  • Models disparate impact using a binary penalty policy where high-risk individuals face incarceration (t_H=1), deriving expected incarceration probability as E[T] = P(T≠0).
  • Establishes a bound on disparate impact Δ using total variation distance d_TV(f_b,y, f_w,y), showing Δ ≤ (t_H - t_L) * d_TV.

Experimental results

Research questions

  • RQ1How does group-specific recidivism prevalence interact with a well-calibrated RPI to produce disparate impact?
  • RQ2To what extent do false positive and false negative rates differ across racial groups in a well-calibrated RPI like COMPAS?
  • RQ3Can a fairness criterion based on calibration still result in unequal treatment outcomes under a binary penalty policy?
  • RQ4Is balancing overall error rates across groups sufficient to ensure fairness at finer levels of subgroup granularity?
  • RQ5What is the relationship between statistical effect size measures (e.g., d_TV) and the magnitude of disparate impact?

Key findings

  • Despite being well-calibrated (i.e., free from predictive bias), the COMPAS RPI produces higher false positive rates for Black defendants (FPR_b ≈ 0.45) than for White defendants (FPR_w ≈ 0.25) due to differing recidivism prevalence.
  • The disparity in false positive rates leads to a 1.8-fold increase in the likelihood of incarceration for non-recidivating Black defendants compared to White defendants under a binary penalty policy.
  • Even within low-prior-record subgroups (e.g., misdemeanor offenders), significant differences in false positive rates persist between Black and White defendants.
  • The total variation distance between Black and White defendant score distributions (d_TV = 24.5%) provides a sharp upper bound on the magnitude of disparate impact.
  • Cohen’s d for the COMPAS score distribution is 0.60, indicating a moderate effect size, but the non-normality of the scores limits the validity of traditional non-overlap measures.
  • The study shows that balancing overall error rates is insufficient to eliminate disparities at finer subgroup levels, such as by prior conviction count.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.