Skip to main content
QUICK REVIEW

[Paper Review] A Hierarchy of Limitations in Machine Learning

Momin M. Malik|arXiv (Cornell University)|Feb 12, 2020
Explainable Artificial Intelligence (XAI)232 references49 citations
TL;DR

A structured, multi-level critique of the fundamental conceptual, procedural, and statistical limitations of machine learning when applied to society, focusing on how downstream consequences propagate through a hierarchy from quantification to cross-validation.

ABSTRACT

"All models are wrong, but some are useful", wrote George E. P. Box (1979). Machine learning has focused on the usefulness of probability models for prediction in social systems, but is only now coming to grips with the ways in which these models are wrong---and the consequences of those shortcomings. This paper attempts a comprehensive, structured overview of the specific conceptual, procedural, and statistical limitations of models in machine learning when applied to society. Machine learning modelers themselves can use the described hierarchy to identify possible failure points and think through how to address them, and consumers of machine learning models can know what to question when confronted with the decision about if, where, and how to apply machine learning. The limitations go from commitments inherent in quantification itself, through to showing how unmodeled dependencies can lead to cross-validation being overly optimistic as a way of assessing model performance.

Motivation & Objective

  • Identify and organize the fundamental assumptions and limitations of machine learning when applied to social systems.
  • Explain how dependencies and measurement choices bias model evaluation, especially through cross-validation.
  • Recommend how mixed methods and alternative approaches can address limitations of a machine-learning–only workflow.
  • Provide guidance for modelers and users on where to question ML use in practice.

Proposed method

  • Proposes a four-decision hierarchy guiding ML use: (1) quantitative over qualitative analysis, (2) probabilistic modeling over other modeling, (3) predictive modeling over explanatory modeling, (4) cross-validation as evaluation tool.
  • Develops the concept of a logical chain of custody showing how limitations propagate from quantification to cross-validation.
  • Introduces an extension of optimizer/optimism ideas (Efron, 2004) to theory of how dependencies bias cross-validation.
  • Draws on philosophical, sociological, statistical, and ML critiques to connect the abstractions of ML with social context.
  • Cites and discusses constructs, latent variables, and the role of ground truth vs. constructs in measurement.

Experimental results

Research questions

  • RQ1What are the fundamental assumptions and limitations of using machine learning for social analysis?
  • RQ2How do dependencies and measurement choices bias cross-validation and model evaluation?
  • RQ3What is the impact of prioritizing quantitative, probabilistic, and predictive approaches on understanding social phenomena?
  • RQ4How can mixed methods address the limitations of ML when applied to society?
  • RQ5What framework can help modelers and consumers interrogate ML claims more effectively?

Key findings

  • There is a hierarchical chain of limitations that propagate through quantification, constructs, and cross-validation.
  • Cross-validation can be overly optimistic about generalizability when dependencies are present and not adequately accounted for.
  • Quantification imposes a central-tendency view and may misrepresent meaning-making and lived experience.
  • Constructs and measurement issues can lead to ground-truth proxies that misrepresent underlying latent factors.
  • Mixed methods and alternative validation approaches can mitigate some ML limitations, though they require careful collaboration and methodological work.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.