[Paper Review] Promises and Challenges of Causality for Ethical Machine Learning
This paper proposes a causal fairness framework based on the potential outcomes model to address limitations in statistical fairness metrics by emphasizing timing and nature of interventions on perceived social categories rather than immutable attributes. It demonstrates through synthetic and real-world police stop data that causal analysis at different decision stages reveals varying fairness violations, with the proposed method achieving zero disparity in key fairness metrics while maintaining high accuracy.
In recent years, there has been increasing interest in causal reasoning for designing fair decision-making systems due to its compatibility with legal frameworks, interpretability for human stakeholders, and robustness to spurious correlations inherent in observational data, among other factors. The recent attention to causal fairness, however, has been accompanied with great skepticism due to practical and epistemological challenges with applying current causal fairness approaches in the literature. Motivated by the long-standing empirical work on causality in econometrics, social sciences, and biomedical sciences, in this paper we lay out the conditions for appropriate application of causal fairness under the "potential outcomes framework." We highlight key aspects of causal inference that are often ignored in the causal fairness literature. In particular, we discuss the importance of specifying the nature and timing of interventions on social categories such as race or gender. Precisely, instead of postulating an intervention on immutable attributes, we propose a shift in focus to their perceptions and discuss the implications for fairness evaluation. We argue that such conceptualization of the intervention is key in evaluating the validity of causal assumptions and conducting sound causal analysis including avoiding post-treatment bias. Subsequently, we illustrate how causality can address the limitations of existing fairness metrics, including those that depend upon statistical correlations. Specifically, we introduce causal variants of common statistical notions of fairness, and we make a novel observation that under the causal framework there is no fundamental disagreement between different notions of fairness. Finally, we conduct extensive experiments where we demonstrate our approach for evaluating and mitigating unfairness, specially when post-treatment variables are present.
Motivation & Objective
- To address the limitations of statistical fairness metrics that rely on passive correlations and conflict across criteria.
- To resolve skepticism around causal fairness by clarifying valid assumptions and intervention timing.
- To shift focus from immutable social categories to their perceptions in causal fairness analysis.
- To mitigate post-treatment bias and improve validity of causal fairness assessments in real-world decision systems.
- To demonstrate that causal fairness criteria are fundamentally compatible when interventions are properly timed and conceptualized.
Proposed method
- Adopt the potential outcomes framework to define counterfactual fairness under hypothetical interventions.
- Distinguish between interventions on immutable attributes (e.g., race at birth) and their perceptions (e.g., perceived race in decision-making contexts).
- Model interventions at different temporal stages (e.g., search vs. arrest in policing) to assess fairness at multiple decision points.
- Use neural networks with two hidden layers to impute post-treatment variables (e.g., search outcome, arrest) under counterfactual conditions.
- Apply causal parity metrics to evaluate fairness violations at each stage, comparing observed and counterfactual outcomes.
- Filter counterfactual outcomes based on intervention conditions (e.g., if no search occurs, set outcome to 'nothing found').
Experimental results
Research questions
- RQ1How can causal fairness be meaningfully applied when sensitive attributes are immutable and cannot be manipulated?
- RQ2What is the impact of intervention timing on causal fairness assessment in multi-stage decision systems?
- RQ3How do perceptions of social categories differ from the attributes themselves in causal fairness modeling?
- RQ4Can causal fairness metrics resolve the inherent conflicts between statistical fairness criteria?
- RQ5What role does post-treatment bias play in misleading fairness evaluations, and how can it be avoided?
Key findings
- Intervening at the search stage revealed a 5.5% disparity in search rates between Black and White individuals, increasing to 8.2% in arrest rates.
- When interventions were delayed to the arrest stage, the disparity in arrest rates dropped to 3.9%, highlighting the importance of timing in causal analysis.
- The proposed causal model achieved zero disparity in ReW (-0.0258), PRem (-0.024), and ROC (-0.028) compared to baseline models.
- The causal approach outperformed statistical fairness metrics across all criteria while maintaining high accuracy (0.768–0.933) on real-world datasets.
- Causal parity violations were significantly reduced when accounting for discriminatory effects at earlier decision stages, such as search.
- The study demonstrated that different fairness criteria are not fundamentally incompatible under a proper causal framework when intervention timing and nature are correctly specified.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.