Skip to main content
QUICK REVIEW

[Paper Review] Bias in Data-driven AI Systems -- An Introductory Survey

Eirini Ntoutsi, Pavlos Fafalios|arXiv (Cornell University)|Jan 14, 2020
Ethics and Social Impacts of AI86 references21 citations
TL;DR

This survey provides a comprehensive, multi-disciplinary overview of bias in data-driven AI systems, categorizing approaches into understanding, mitigating, and accounting for bias. It integrates technical solutions with legal and ethical frameworks, emphasizing the need for fairness-aware AI design through pre-processing, in-processing, and post-processing techniques, while highlighting the limitations of purely technical fixes and the necessity of legal and societal collaboration.

ABSTRACT

AI-based systems are widely employed nowadays to make decisions that have far-reaching impacts on individuals and society. Their decisions might affect everyone, everywhere and anytime, entailing concerns about potential human rights issues. Therefore, it is necessary to move beyond traditional AI algorithms optimized for predictive performance and embed ethical and legal principles in their design, training and deployment to ensure social good while still benefiting from the huge potential of the AI technology. The goal of this survey is to provide a broad multi-disciplinary overview of the area of bias in AI systems, focusing on technical challenges and solutions as well as to suggest new research directions towards approaches well-grounded in a legal frame. In this survey, we focus on data-driven AI, as a large part of AI is powered nowadays by (big) data and powerful Machine Learning (ML) algorithms. If otherwise not specified, we use the general term bias to describe problems related to the gathering or processing of data that might result in prejudiced decisions on the bases of demographic features like race, sex, etc.

Motivation & Objective

  • To provide a broad, multi-disciplinary overview of bias in data-driven AI systems, integrating technical, ethical, and legal perspectives.
  • To identify and categorize technical approaches for understanding, mitigating, and accounting for bias in AI decision-making.
  • To examine the legal foundations of AI fairness, particularly within the EU regulatory context, and highlight gaps in current legislation.
  • To emphasize the limitations of technical solutions alone and advocate for interdisciplinary collaboration between technologists, legal experts, and society.
  • To guide future research toward ethically grounded, legally compliant, and socially responsible AI systems.

Proposed method

  • Categorizes bias mitigation into three stages: pre-processing (bias correction in training data), in-processing (fairness-aware learning algorithms), and post-processing (adjustment of model outputs).
  • Reviews formal definitions of fairness, including demographic parity, equal opportunity, and equalized odds, to clarify technical and ethical trade-offs.
  • Analyzes case studies such as COMPAS (racial bias in recidivism prediction) and Google Ads (gender bias in job advertising) to illustrate real-world manifestations.
  • Examines emerging techniques like GAN-based synthetic data generation for fairness, while cautioning about representativeness and bias amplification risks.
  • Proposes strategies to reduce cognitive bias in development teams, such as increasing diversity and enabling algorithmic transparency through reverse engineering.
  • Reviews legal and regulatory developments, including the EU’s Ethics Guidelines for Trustworthy AI and ISO 8000 on data quality, to assess their relevance to fairness in AI.

Experimental results

Research questions

  • RQ1How does bias enter data-driven AI systems, and what are the primary sources of unfairness in training data and model predictions?
  • RQ2What technical methods exist for detecting, measuring, and mitigating bias at different stages of the AI lifecycle (pre-, in-, and post-processing)?
  • RQ3How can fairness in AI be legally and ethically grounded, and what role do regulations and standards play in shaping responsible AI design?
  • RQ4To what extent can technical solutions alone address systemic societal biases, and what are the limitations of current fairness-aware machine learning approaches?
  • RQ5What are the key challenges in ensuring representativeness and fairness in synthetic data generation using GANs and other generative models?

Key findings

  • Bias in AI systems often stems from historical and societal inequities, which are amplified by data collection and model training processes, as seen in the COMPAS and Google Ads cases.
  • Technical approaches to fairness—such as pre-processing reweighting, in-processing adversarial debiasing, and post-processing threshold adjustment—can reduce demographic disparities but involve trade-offs between fairness and predictive performance.
  • The use of Generative Adversarial Networks (GANs) to create synthetic fair data is promising but risks perpetuating or worsening representativeness issues if training data is biased.
  • Cognitive biases in development teams, particularly the lack of awareness among privileged groups about identity categories, contribute significantly to the emergence of bias in AI systems.
  • Current legal frameworks in the EU, such as the Ethics Guidelines for Trustworthy AI and the Draft AI Act, remain generic and lack binding standards for algorithmic fairness, highlighting a need for more specific regulation.
  • There is no consensus on algorithmic fairness regulations across EU member states, and effective legislation will require close collaboration between legal experts and technical researchers.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.