[Paper Review] Group Fairness: Independence Revisited
This paper re-evaluates independence (statistical parity) as a measure of group fairness in algorithmic decision-making, countering recent criticisms by showing that arguments against it are either equally applicable to other fairness measures like separation and sufficiency, or based on flawed assumptions. It demonstrates that independence captures a distinct form of fairness not addressed by other criteria and advocates for its inclusion in fairness evaluations alongside other measures.
This paper critically examines arguments against independence, a measure of group fairness also known as statistical parity and as demographic parity. In recent discussions of fairness in computer science, some have maintained that independence is not a suitable measure of group fairness. This position is at least partially based on two influential papers (Dwork et al., 2012, Hardt et al., 2016) that provide arguments against independence. We revisit these arguments, and we find that the case against independence is rather weak. We also give arguments in favor of independence, showing that it plays a distinctive role in considerations of fairness. Finally, we discuss how to balance different fairness considerations.
Motivation & Objective
- To critically reassess the validity of arguments against independence (statistical parity) as a fairness measure in machine learning.
- To challenge the claim that independence is inherently flawed or less defensible than separation or sufficiency in fairness evaluations.
- To demonstrate that independence highlights a distinct kind of unfairness not captured by other fairness measures.
- To show that the conservative nature of fairness measures—particularly sufficiency and separation—is not preserved under increased accuracy, undermining a key argument against independence.
- To advocate for a balanced, context-sensitive approach to fairness that includes independence alongside other fairness criteria.
Proposed method
- Re-analyzes foundational arguments against independence from Dwork et al. (2012) and Hardt et al. (2016), showing they are not uniquely applicable to independence.
- Introduces the concept of 'incremental conservativeness' to evaluate whether fairness measures are preserved when predictor accuracy improves.
- Constructs counterexamples using confusion matrices to show that increasing accuracy can violate both separation and sufficiency, proving they are not incrementally conservative.
- Uses conditional independence formalism (e.g., R ⊥ A, Y ⊥ A|R, R ⊥ A|Y) to define and compare fairness measures in binary classification settings.
- Applies the framework to real-world examples like college admissions and criminal risk assessment to illustrate context-dependent fairness trade-offs.
- Employs probabilistic reasoning and proof techniques to establish that independence is not logically incompatible with high accuracy when group base rates are unequal.
Experimental results
Research questions
- RQ1Are the arguments against independence in fairness metrics uniquely applicable to independence, or do they equally undermine other fairness measures?
- RQ2Is the claim that independence fails to be 'conservative' (i.e., preserved under accuracy gains) valid, and how does it compare to sufficiency and separation?
- RQ3Does independence capture a distinct form of fairness not addressed by separation or sufficiency?
- RQ4How should multiple fairness criteria—especially independence, separation, and sufficiency—be balanced in practice?
- RQ5What role should context-specific considerations (e.g., cost of errors) play in selecting fairness measures?
Key findings
- The arguments against independence from Dwork et al. (2012) and Hardt et al. (2016) are not uniquely applicable to independence, as they equally challenge other fairness measures like separation and sufficiency.
- Sufficiency and separation are not incrementally conservative: increasing predictor accuracy can violate both measures, even when the original predictor satisfied them.
- Independence captures a distinct fairness dimension—group-based prediction balance—that is not subsumed by separation or sufficiency, particularly in cases where group base rates differ.
- Independence is compatible with high accuracy when the true label is not independent of group membership, contrary to claims that it inherently conflicts with accuracy.
- The paper demonstrates via counterexamples that increasing accuracy by shifting false negatives to true positives (non-proportionally) can break both separation and sufficiency, even while improving overall performance.
- The study concludes that independence should be considered a legitimate and necessary component of fairness evaluation, especially when other measures conflict or fail to capture group-level imbalance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.