[Paper Review] Fairness in Machine Learning: A Survey
This survey provides an entry-level overview of fairness in ML, organizing approaches into pre-processing, in-processing, and post-processing, and detailing metrics, methods, toolkits, and open challenges.
As Machine Learning technologies become increasingly used in contexts that affect citizens, companies as well as researchers need to be confident that their application of these methods will not have unexpected social implications, such as bias towards gender, ethnicity, and/or people with disabilities. There is significant literature on approaches to mitigate bias and promote fairness, yet the area is complex and hard to penetrate for newcomers to the domain. This article seeks to provide an overview of the different schools of thought and approaches to mitigating (social) biases and increase fairness in the Machine Learning literature. It organises approaches into the widely accepted framework of pre-processing, in-processing, and post-processing methods, subcategorizing into a further 11 method areas. Although much of the literature emphasizes binary classification, a discussion of fairness in regression, recommender systems, unsupervised learning, and natural language processing is also provided along with a selection of currently available open source libraries. The article concludes by summarising open challenges articulated as four dilemmas for fairness research.
Motivation & Objective
- Introduce readers to the key concepts and history of fairness in ML.
- Summarize and standardize fairness metrics and their trade-offs.
- Provide a two-dimensional taxonomy of fair ML approaches (pre-processing, in-processing, post-processing) and extend to non-binary tasks.
- Highlight commonly used toolkits and practical considerations, including legal and social accountability.
- Identify open challenges and four dilemmas guiding future fairness research.
Proposed method
- Classifies fairness techniques into pre-processing, in-processing, and post-processing within a unified intervention framework.
- Catalogues and contrasts a wide range of fairness metrics (group, individual, and counterfactual) with a shared notation.
- Discusses abstract fairness criteria (Independence, Separation, Sufficiency) and their impossibility results.
- Reviews method families (blinding, causal methods, sampling/subgroup analysis) and their eligibility for different ML stages.
- Outlines practical considerations such as data proxies, protected variables, and potential legal/interpretability implications.
- Summarizes available open-source libraries and maps current research directions to four future dilemmas.
Experimental results
Research questions
- RQ1What are the main methodological categories for achieving fairness in ML, and how do they relate to data and model stages (pre-, in-, post-processing)?
- RQ2How can different fairness metrics be defined, interpreted, and traded off against accuracy in binary and non-binary tasks?
- RQ3What are the forces and limitations of causal, blinding, and sampling approaches in achieving fair outcomes?
- RQ4What practical tools and open challenges exist for deploying fair ML in real-world settings?
Key findings
- There is no universal definition of fairness; multiple metrics capture different notions (statistical parity, equalized odds, calibration, etc.) with inherent trade-offs.
- Pre-processing, in-processing, and post-processing provide flexible but not universally comparable intervention points, each with unique interpretability and legal implications.
- The literature highlights a tension between group fairness and individual fairness, with many impossibility results showing incompatible objectives under certain conditions.
- Causal, proxy-identifier, and graph-based methods help identify biases and proxies but require substantial background information and can be computationally intensive.
- A widening array of open-source libraries supports fair ML, yet practical adoption remains challenged by data quality, proxy variables, and dynamic data shifts.
- Researchers emphasize four dilemmas guiding future work, focusing on accessibility, accountability, and societal impact.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.