Skip to main content
QUICK REVIEW

[Paper Review] Wrong side of the tracks: Big Data and Protected Categories

Simon DeDeo|arXiv (Cornell University)|Dec 15, 2014
Adversarial Robustness in Machine Learning15 references11 citations
TL;DR

This paper addresses the ethical challenge of algorithmic discrimination in Big Data by showing how machine learning models can inadvertently make decisions based on protected categories (e.g., race, sex) due to strong correlations in data. It proposes using information theory to decorrelate predictions from protected variables with minimal accuracy loss, while advocating for future causal-aware algorithms to ensure ethical transparency and democratic accountability in automated decision-making.

ABSTRACT

When we use machine learning for public policy, we find that many useful variables are associated with others on which it would be ethically problematic to base decisions. This problem becomes particularly acute in the Big Data era, when predictions are often made in the absence of strong theories for underlying causal mechanisms. We describe the dangers to democratic decision-making when high-performance algorithms fail to provide an explicit account of causation. We then demonstrate how information theory allows us to degrade predictions so that they decorrelate from protected variables with minimal loss of accuracy. Enforcing total decorrelation is at best a near-term solution, however. The role of causal argument in ethical debate urges the development of new, interpretable machine-learning algorithms that reference causal mechanisms.

Motivation & Objective

  • To identify how Big Data algorithms can produce ethically problematic decisions by relying on correlations with protected categories such as race, sex, or socioeconomic status.
  • To demonstrate that even accurate, well-intentioned algorithms can perpetuate or amplify discrimination when they lack causal transparency.
  • To propose a near-term technical solution—information-theoretic decorrelation—that reduces reliance on protected variables without significant loss in prediction accuracy.
  • To advocate for the development of interpretable, causal machine learning models that can support ethical and democratic decision-making in policy and governance.
  • To emphasize the need for algorithmic transparency and public debate by 'opening the algorithmic box' to enable scrutiny of hidden assumptions and moral reasoning in automated systems.

Proposed method

  • Uses information theory to mathematically degrade predictions so they become decorrelated from protected variables, minimizing accuracy loss.
  • Applies the concept of 'outcome equalization' as a limiting case to assess the impact of removing correlations with protected categories.
  • Proposes contribution propagation and the Bayesian list machine as emerging techniques to reverse-engineer implicit causal models in algorithms.
  • Employs a framework of causal reasoning to evaluate whether algorithmic decisions align with ethical norms, especially in contexts like sentencing or lending.
  • Introduces a method to enforce statistical independence between predictions and protected attributes while preserving predictive performance.
  • Calls for institutional commitments from corporations and governments to disclose key features of their algorithms to enable public and expert scrutiny.

Experimental results

Research questions

  • RQ1How can machine learning models inadvertently make decisions based on protected categories due to data correlations, even without explicit use of those variables?
  • RQ2To what extent can predictive accuracy be preserved while decorrelating model outputs from protected variables like race or sex?
  • RQ3What mathematical and computational methods can be used to reverse-engineer the implicit causal models embedded in opaque algorithms?
  • RQ4How does the absence of causal explanation in high-performance algorithms undermine democratic deliberation and ethical accountability?
  • RQ5In what ways can information-theoretic approaches serve as a bridge toward more transparent and ethically sound algorithmic decision-making?

Key findings

  • Strong correlations between protected categories (e.g., race, sex) and non-protected variables (e.g., shopping habits, electricity use) can lead to indirect discrimination even when those categories are not explicitly used in models.
  • Information-theoretic decorrelation allows predictions to be made independent of protected variables with minimal loss in accuracy, offering a practical near-term solution to algorithmic bias.
  • The outcome equalization method serves as a limiting case that reveals how much a model's predictions depend on protected attributes, highlighting problematic correlations.
  • Current high-performance algorithms often fail to provide causal explanations, making it difficult to assess whether decisions are ethically justifiable or merely statistically convenient.
  • Without transparency, democratic debate on fairness and equity becomes impossible, as citizens and policymakers cannot scrutinize the models shaping public policy.
  • The development of interpretable, causal machine learning models—such as contribution propagation and the Bayesian list machine—offers a path toward ethical, transparent, and accountable algorithmic systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.