Skip to main content
QUICK REVIEW

[Paper Review] Bias in Machine Learning -- What is it Good for?

Thomas Hellström, Virginia Dignum|arXiv (Cornell University)|Apr 1, 2020
Ethics and Social Impacts of AI40 references46 citations
TL;DR

This paper proposes a taxonomy of the various meanings of bias in machine learning, distinguishes biases along the learning pipeline, and discusses their interrelations and implications for model bias and societal fairness.

ABSTRACT

In public media as well as in scientific publications, the term \emph{bias} is used in conjunction with machine learning in many different contexts, and with many different meanings. This paper proposes a taxonomy of these different meanings, terminology, and definitions by surveying the, primarily scientific, literature on machine learning. In some cases, we suggest extensions and modifications to promote a clear terminology and completeness. The survey is followed by an analysis and discussion on how different types of biases are connected and depend on each other. We conclude that there is a complex relation between bias occurring in the machine learning pipeline that leads to a model, and the eventual bias of the model (which is typically related to social discrimination). The former bias may or may not influence the latter, in a sometimes bad, and sometime good way.

Motivation & Objective

  • Clarify the different usages and definitions of bias in machine learning across literature.
  • Present a taxonomy of biases along the machine learning pipeline (world, data generation, learning).
  • Discuss how different bias types interact and influence the final model bias, including ethical and causal considerations.

Proposed method

  • Survey of published research to categorize and define biases encountered in ML.
  • Introduce and standardize terminology through descriptive names for shared concepts.
  • Propose a taxonomy diagram and discuss connections between bias types.
  • Discuss causal considerations and the distinction between world-as-it-is versus world-as-it-should-be.

Experimental results

Research questions

  • RQ1What are the distinct notions of bias used in relation to machine learning across literature?
  • RQ2How do biases arising in the world, data generation, and learning stages relate to the bias observed in final models?
  • RQ3What terminology and extensions are needed to achieve a clear and complete bias taxonomy in ML?

Key findings

  • There are multiple, sometimes conflicting, notions of bias in ML that span learning, data, and world-related factors.
  • A taxonomy linking historical/world bias, data-generation bias, and learning bias helps explain how biases propagate to model bias.
  • Model bias is influenced by causal factors and may be desirable or undesirable depending on the task and normative goals.
  • Many bias notions are interrelated and cannot be avoided simultaneously due to trade-offs in classifier performance metrics.
  • Debiasing can target either the world as it should be or the data used to train models, each with different implications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.