Skip to main content
QUICK REVIEW

[Paper Review] Designing Algorithmic Recommendations to Achieve Human-AI Complementarity

Bryce McLaughlin, Jann Spiess|arXiv (Cornell University)|May 2, 2024
Ethics and Social Impacts of AI4 citations
TL;DR

This paper proposes a causal inference-based framework for designing algorithmic recommendation systems that enhance human-AI complementarity in binary decision-making. By modeling human compliance and response to recommendations using potential outcomes and monotonicity assumptions, the authors show that algorithms designed to provide complementary information—rather than direct predictions—significantly improve human decision performance in a controlled hiring experiment.

ABSTRACT

Algorithms frequently assist, rather than replace, human decision-makers. However, the design and analysis of algorithms often focus on predicting outcomes and do not explicitly model their effect on human decisions. This discrepancy between the design and role of algorithmic assistants becomes particularly concerning in light of empirical evidence that suggests that algorithmic assistants again and again fail to improve human decisions. In this article, we formalize the design of recommendation algorithms that assist human decision-makers without making restrictive ex-ante assumptions about how recommendations affect decisions. We formulate an algorithmic-design problem that leverages the potential-outcomes framework from causal inference to model the effect of recommendations on a human decision-maker's binary treatment choice. Within this model, we introduce a monotonicity assumption that leads to an intuitive classification of human responses to the algorithm. Under this assumption, we can express the human's response to algorithmic recommendations in terms of their compliance with the algorithm and the active decision they would take if the algorithm sends no recommendation. We showcase the utility of our framework using an online experiment that simulates a hiring task. We argue that our approach can make sense of the relative performance of different recommendation algorithms in the experiment and can help design solutions that realize human-AI complementarity. Finally, we leverage our approach to derive minimax optimal recommendation algorithms that can be implemented with machine learning using limited training data.

Motivation & Objective

  • To address the gap in algorithm design that often ignores how recommendations actually affect human decision-making, especially when human judgment remains final.
  • To formalize the design of recommendation algorithms in a principal–agent model where the human retains decision authority, enabling complementarity between human and algorithmic expertise.
  • To develop a framework that explicitly models human response to recommendations—particularly compliance and active decision-making—using causal inference principles.
  • To demonstrate empirically that recommendation algorithms designed for complementarity outperform those focused solely on predictive accuracy in human-assisted decision tasks.
  • To provide a principled approach to optimizing recommendation systems by decomposing performance into algorithmic accuracy and human response effects.

Proposed method

  • Formalizes the recommendation design problem using the potential-outcomes framework from causal inference, modeling outcomes under different recommendation conditions.
  • Introduces a monotonicity assumption that restricts human responses to recommendations such that recommendations only move decisions toward the recommended action.
  • Decomposes the objective function into two components: (1) the loss from directly implementing algorithmic recommendations and (2) the loss from human deviation from recommendations.
  • Uses the instrumental variable analogy from causal inference to model compliance behavior, treating recommendations as instruments that affect human decisions.
  • Applies the framework to an online experiment simulating a hiring task with 961 subjects, varying recommendation algorithms based on different assumptions about compliance.
  • Empirically evaluates performance across treatment arms (e.g., Predictive, Complementary, Triage) and uses comprehension checks to validate that understanding of the algorithm affects outcomes.

Experimental results

Research questions

  • RQ1How can recommendation algorithms be designed to improve human decision-making rather than merely predicting outcomes?
  • RQ2What role does human compliance with algorithmic recommendations play in determining overall system performance?
  • RQ3Under what conditions does human–AI complementarity emerge in binary decision tasks?
  • RQ4How do different algorithmic designs (e.g., predictive vs. complementary) affect human performance in real-world decision scenarios?
  • RQ5To what extent does a human decision-maker’s understanding of the recommendation algorithm influence the success of human–AI collaboration?

Key findings

  • Subjects performed significantly better with the Complementary and Complementary Triage recommendation algorithms than with the Control (no algorithm) or Predictive algorithm alone, achieving human–AI complementarity.
  • The Complementary algorithm, which provides information that supports but does not dictate decisions, outperformed all other tested algorithms in the experiment.
  • Subjects in the Predictive treatment matched the performance of the algorithm alone or the human alone, failing to achieve any improvement through collaboration.
  • Subjects who correctly understood the recommendation algorithm (as measured by comprehension questions) showed stronger, more consistent performance, supporting the importance of algorithmic transparency.
  • Compliance behavior was generally sophisticated—subjects overrode recommendations for Engineering candidates when appropriate—indicating that humans can make nuanced, context-aware decisions when given complementary support.
  • The framework successfully explained the relative performance of different algorithms, validating its utility in predicting and designing effective human–AI collaboration systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.