Skip to main content
QUICK REVIEW

[Paper Review] Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes

Connor Toups, Rishi Bommasani|arXiv (Cornell University)|Jul 12, 2023
COVID-19 epidemiological studies4 citations
TL;DR

This paper introduces ecosystem-level analysis to study the societal impact of deployed machine learning by examining collective model outcomes rather than individual models. It reveals that systemic failures—where users are misclassified by all models—are widespread and persist despite model improvements, with racial disparities in dermatology models becoming more pronounced under this framework, highlighting limitations of traditional fairness metrics.

ABSTRACT

Machine learning is traditionally studied at the model level: researchers measure and improve the accuracy, robustness, bias, efficiency, and other dimensions of specific models. In practice, the societal impact of machine learning is determined by the surrounding context of machine learning deployments. To capture this, we introduce ecosystem-level analysis: rather than analyzing a single model, we consider the collection of models that are deployed in a given context. For example, ecosystem-level analysis in hiring recognizes that a job candidate's outcomes are not only determined by a single hiring algorithm or firm but instead by the collective decisions of all the firms they applied to. Across three modalities (text, images, speech) and 11 datasets, we establish a clear trend: deployed machine learning is prone to systemic failure, meaning some users are exclusively misclassified by all models available. Even when individual models improve at the population level over time, we find these improvements rarely reduce the prevalence of systemic failure. Instead, the benefits of these improvements predominantly accrue to individuals who are already correctly classified by other models. In light of these trends, we consider medical imaging for dermatology where the costs of systemic failure are especially high. While traditional analyses reveal racial performance disparities for both models and humans, ecosystem-level analysis reveals new forms of racial disparity in model predictions that do not present in human predictions. These examples demonstrate ecosystem-level analysis has unique strengths for characterizing the societal impact of machine learning.

Motivation & Objective

  • To address the limitations of model-level fairness analysis by shifting focus to the collective impact of multiple deployed models on individuals.
  • To investigate whether improvements in individual models reduce systemic failures—cases where all models misclassify a user.
  • To uncover new forms of racial disparity in machine learning outcomes that are invisible in traditional group-level fairness analyses.
  • To evaluate the impact of model improvements on marginalized individuals who are systematically failed by all models in the ecosystem.
  • To demonstrate that ecosystem-level analysis reveals more nuanced and consequential societal impacts than conventional model-centric evaluation.

Proposed method

  • Define the failure matrix F as the collection of outcomes individuals receive from all decision-makers in a given ecosystem.
  • Identify systemic failures as individuals who receive negative outcomes from every model in the ecosystem.
  • Use the HAPI audit framework to analyze 11 datasets across text, speech, and vision modalities with 3 commercial systems per modality.
  • Measure both gross and net improvements in model performance to assess whether progress reduces systemic failures.
  • Conduct ecosystem-level analysis on dermatology models and human dermatologists using the DDI dataset with Fitzpatrick skin type annotations.
  • Compare model and human outcomes using profile polarization and racial disparity metrics, excluding HAM10k due to its near-universal negative predictions.
Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes

Experimental results

Research questions

  • RQ1To what extent do systemic failures—where all models misclassify a user—persist across deployed machine learning systems?
  • RQ2Do improvements in individual models lead to reductions in systemic failures, or are benefits concentrated among users already correctly classified by other models?
  • RQ3How do racial disparities in model performance manifest under ecosystem-level analysis compared to traditional fairness metrics?
  • RQ4Do dermatology models exhibit greater polarization and racial bias in their collective outcomes than human dermatologists?
  • RQ5How do net improvements in model performance compare to gross improvements in terms of progress on systemic failures?

Key findings

  • Systemic failures—where all models misclassify a user—are prevalent across 11 datasets spanning text, speech, and vision, with polarization rates significantly exceeding independent model behavior predictions.
  • Despite a 2.5% reduction in error rate on the waimai dataset, Amazon’s sentiment analysis API made zero gross improvements on instances systematically failed by all other models.
  • On average, only 10% of a model’s instance-level improvements occur on cases misclassified by all other models, even though systemic failures account for 27% of all cases that could be improved.
  • In dermatology, models show increased polarization for darker skin tones, a disparity not present in human predictions, revealing new forms of racial inequity invisible to traditional fairness analysis.
  • Including the HAM10k model exacerbates profile polarization and racial disparities, confirming that findings are robust and not artifacts of model exclusion.
  • Net improvements show even less progress on systemic failures than gross improvements, indicating that model gains are largely concentrated on already-well-served individuals.
Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.