Skip to main content
QUICK REVIEW

[Paper Review] "If it didn't happen, why would I change my decision?": How Judges Respond to Counterfactual Explanations for the Public Safety Assessment

Yaniv Yacoby, Ben Green|arXiv (Cornell University)|May 11, 2022
Explainable Artificial Intelligence (XAI)4 citations
TL;DR

This study investigates how U.S. state court judges interpret and respond to counterfactual explanations (CFEs) in the Public Safety Assessment (PSA), a pretrial risk algorithm. Despite being trained to understand hypotheticals, judges initially misread CFEs as factual changes and later ignored them, asserting their role was to decide only on actual defendant circumstances—highlighting a fundamental disconnect between CFE design and judicial reasoning workflows.

ABSTRACT

Many researchers and policymakers have expressed excitement about algorithmic explanations enabling more fair and responsible decision-making. However, recent experimental studies have found that explanations do not always improve human use of algorithmic advice. In this study, we shed light on how people interpret and respond to counterfactual explanations (CFEs) -- explanations that show how a model's output would change with marginal changes to its input(s) -- in the context of pretrial risk assessment instruments (PRAIs). We ran think-aloud trials with eight sitting U.S. state court judges, providing them with recommendations from a PRAI that includes CFEs. We found that the CFEs did not alter the judges' decisions. At first, judges misinterpreted the counterfactuals as real -- rather than hypothetical -- changes to defendants. Once judges understood what the counterfactuals meant, they ignored them, stating their role is only to make decisions regarding the actual defendant in question. The judges also expressed a mix of reasons for ignoring or following the advice of the PRAI without CFEs. These results add to the literature detailing the unexpected ways in which people respond to algorithms and explanations. They also highlight new challenges associated with improving human-algorithm collaborations through explanations.

Motivation & Objective

  • To examine how sitting U.S. state court judges interpret and respond to counterfactual explanations (CFEs) in the context of the Public Safety Assessment (PSA), a pretrial risk assessment tool.
  • To investigate whether CFEs—showing how risk scores would change under hypothetical input variations—alter judges’ decision-making or understanding of algorithmic recommendations.
  • To explore the reasons behind judges’ rejection or misunderstanding of CFEs, especially given their legal training in counterfactual reasoning.
  • To assess whether CFEs improve transparency, trust, or critical evaluation of algorithmic advice in judicial decision-making.

Proposed method

  • Conducted think-aloud cognitive protocol studies with eight sitting U.S. state court judges from two states using the PSA.
  • Presented judges with hypothetical pretrial cases, including PSA risk scores and CFEs showing how risk scores would change under minor input modifications (e.g., removing a prior felony conviction).
  • Used two rounds of testing: Round 1 for initial interpretation, Round 2 with additional training to clarify CFEs as hypothetical, not factual.
  • Collected verbal protocols to analyze real-time reasoning, focusing on how judges processed CFEs and integrated them into decisions.
  • Employed qualitative analysis to identify patterns in misinterpretation, rejection, and reasoning about algorithmic advice.
  • Explored the impact of prior experience with the PSA-DMF (Decision-Making Framework) on judges’ responses to CFEs.

Experimental results

Research questions

  • RQ1How do judges initially interpret counterfactual explanations (CFEs) in the context of the Public Safety Assessment (PSA)?
  • RQ2To what extent do CFEs influence judges’ risk assessments or decisions about pretrial release?
  • RQ3Why do judges reject or ignore CFEs despite their legal training in counterfactual reasoning?
  • RQ4How does prior experience with the PSA-DMF affect judges’ comprehension and use of CFEs?
  • RQ5What design or instructional approaches could make CFEs more interpretable and actionable for judicial decision-makers?

Key findings

  • Judges initially misinterpreted CFEs as factual changes to defendants’ profiles, not hypotheticals, despite clear labeling.
  • After understanding CFEs as hypothetical, all eight judges consistently ignored them, stating their role was to decide only on the actual defendant before them.
  • Judges expressed no interest in using CFEs to assess model sensitivity or robustness, even after explicit coaching on their potential utility.
  • The judges’ rejection of CFEs persisted even after additional training, suggesting a deeper conflict between CFE logic and judicial reasoning norms.
  • Judges cited concerns about autonomy and professional status as motivations to downplay algorithmic influence, potentially undermining transparency.
  • The study reveals a fundamental mismatch between the counterfactual logic embedded in CFEs and the counterfactual reasoning judges apply to actual case facts, limiting CFEs’ utility in judicial settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.