[Paper Review] Fairness Testing: A Comprehensive Survey and Analysis of Trends
This paper surveys 100 fairness testing studies in ML software, organizing them by testing workflow and components, analyzes trends, datasets, and tools, and highlights research opportunities.
Unfair behaviors of Machine Learning (ML) software have garnered increasing attention and concern among software engineers. To tackle this issue, extensive research has been dedicated to conducting fairness testing of ML software, and this paper offers a comprehensive survey of existing studies in this field. We collect 100 papers and organize them based on the testing workflow (i.e., how to test) and testing components (i.e., what to test). Furthermore, we analyze the research focus, trends, and promising directions in the realm of fairness testing. We also identify widely-adopted datasets and open-source tools for fairness testing.
Motivation & Objective
- Define fairness testing and fairness bugs within ML software from a software engineering perspective.
- Categorize fairness testing research by testing workflow (how to test) and testing components (what to test).
- Summarize publicly available datasets and open-source tools used in fairness testing.
- Analyze trends and identify promising directions and open research opportunities in fairness testing.
Proposed method
- Collected 100 fairness testing papers from DBLP using iterative keyword search and snowballing (backward and forward).
- Expanded collection with author feedback to reach 100 papers.
- Applied thematic synthesis to extract and organize findings around workflow and components.
- Defined fairness bug and fairness testing for ML software in SE terms.
- Provided an overview of datasets and open-source tools used in fairness testing.
Experimental results
Research questions
- RQ1What are the prevailing definitions of fairness used in fairness testing for ML software?
- RQ2How is fairness testing conducted in practice (testing workflow) and what components are tested?
- RQ3What datasets and tools are commonly used for fairness testing?
- RQ4What trends emerge in fairness testing research, and what are the identified opportunities and challenges?
Key findings
- The survey covers 100 papers and synthesizes them into a coherent view of fairness testing in ML software.
- Fairness testing literature is organized around testing workflow (how to test) and testing components (where/what to test).
- There is growing research activity with 89% of fairness testing publications appearing since 2019.
- The paper provides an overview of publicly available datasets and open-source tools for fairness testing.
- A systematic analysis identifies research trends and promising directions for future work in fairness testing.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.