[Paper Review] Multi-dimensional discrimination in Law and Machine Learning -- A comparative overview
This paper provides a comparative overview of multi-dimensional discrimination in law and machine learning, identifying key fairness definitions—especially intersectional fairness—from legal scholarship and analyzing their operationalization in fairness-aware machine learning. It highlights gaps in current ML research, such as limited attention to additive and sequential discrimination, data scarcity for intersectional subgroups, and the need for interdisciplinary alignment between legal and technical fairness frameworks.
AI-driven decision-making can lead to discrimination against certain individuals or social groups based on protected characteristics/attributes such as race, gender, or age. The domain of fairness-aware machine learning focuses on methods and algorithms for understanding, mitigating, and accounting for bias in AI/ML models. Still, thus far, the vast majority of the proposed methods assess fairness based on a single protected attribute, e.g. only gender or race. In reality, though, human identities are multi-dimensional, and discrimination can occur based on more than one protected characteristic, leading to the so-called ``multi-dimensional discrimination'' or ``multi-dimensional fairness'' problem. While well-elaborated in legal literature, the multi-dimensionality of discrimination is less explored in the machine learning community. Recent approaches in this direction mainly follow the so-called intersectional fairness definition from the legal domain, whereas other notions like additive and sequential discrimination are less studied or not considered thus far. In this work, we overview the different definitions of multi-dimensional discrimination/fairness in the legal domain as well as how they have been transferred/ operationalized (if) in the fairness-aware machine learning domain. By juxtaposing these two domains, we draw the connections, identify the limitations, and point out open research directions.
Motivation & Objective
- To examine how multi-dimensional discrimination—based on multiple protected attributes like race, gender, and age—is conceptualized in legal scholarship and operationalized in fairness-aware machine learning.
- To identify the limitations of current fairness-aware ML methods, which predominantly focus on single-attribute fairness, and to highlight the underexplored concepts of additive and sequential discrimination.
- To address the challenge of data scarcity for intersectional subgroups, especially when protected attributes intersect, leading to small or empty subpopulations in training data.
- To bridge the gap between legal and technical fairness by analyzing the transferability of legal fairness concepts into ML frameworks, including challenges in defining and labeling protected attributes.
- To propose open research directions, such as developing statistical tests for small minority groups, improving fairness trade-offs across subgroups, and enabling ethical use of sensitive data under regulations like GDPR and the EU AI Act.
Proposed method
- Systematically reviews legal literature on multi-dimensional discrimination, including intersectionality, cumulative discrimination, and sequential discrimination.
- Maps legal fairness concepts to their operationalization in fairness-aware machine learning, focusing on how intersectional fairness has been adopted and other forms like additive and sequential fairness remain underdeveloped.
- Analyzes the technical challenges in detecting and mitigating multi-dimensional discrimination, particularly data scarcity when multiple protected attributes are combined.
- Evaluates existing fairness metrics and mitigation techniques in ML, identifying their limitations in handling intersectional and sequential discrimination.
- Proposes the use of synthetic data generation as a potential solution to data scarcity, while cautioning about the risk of introducing subgroup biases.
- Examines regulatory frameworks such as GDPR and the EU AI Act, particularly Article 10(5), which permits use of sensitive data for bias detection under strict conditions, to inform ethical data use in fairness research.
Experimental results
Research questions
- RQ1How is multi-dimensional discrimination conceptualized in legal scholarship, and which of these concepts have been operationalized in fairness-aware machine learning?
- RQ2Why has the machine learning community predominantly focused on intersectional fairness while largely neglecting additive and sequential discrimination?
- RQ3What are the technical and ethical challenges in detecting and mitigating multi-dimensional discrimination, especially when protected subgroups become too small for reliable statistical analysis?
- RQ4How can fairness metrics and mitigation techniques be adapted to handle trade-offs between fairness across multiple protected attributes?
- RQ5What role can regulatory frameworks like GDPR and the EU AI Act play in enabling responsible data use for detecting and correcting multi-dimensional discrimination in AI systems?
Key findings
- The vast majority of fairness-aware machine learning methods still focus on mono-dimensional fairness, assessing bias based on a single protected attribute such as gender or race, despite real-world discrimination being inherently multi-dimensional.
- Intersectional fairness—rooted in critical legal theory and the concept of overlapping systems of oppression—has been the primary framework adopted in ML research, but additive and sequential discrimination remain underexplored.
- Data scarcity is a major obstacle: as the number of protected attributes increases, the size of intersectional subgroups shrinks, making statistical analysis and bias detection unreliable, especially for small or marginalized groups.
- Legal jurisprudence, particularly from the European Court of Justice, relies on statistical tests that are often ill-suited for detecting discrimination in small subgroups, highlighting a gap in current fairness evaluation methods.
- The EU AI Act’s Article 10(5), which allows the use of sensitive personal data for bias monitoring, represents a potential regulatory pathway to address data scarcity, though ethical and privacy concerns remain.
- Synthetic data generation is a promising but risky approach to augment rare subgroups, as it may embed or amplify existing biases if not carefully designed with input from social science and legal expertise.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.