Skip to main content
QUICK REVIEW

[Paper Review] Human-AI Collaboration in Decision-Making: Beyond Learning to Defer

Diogo Leitão, Pedro Saleiro|arXiv (Cornell University)|Jun 27, 2022
Human-Automation Interaction and Safety20 citations
TL;DR

The paper analyzes Learning to Defer (L2D) in human–AI decision-making, identifies its key limitations for real-world deployment, and outlines research directions to build more robust, fair, and dynamic HAIC systems that go beyond L2D.

ABSTRACT

Human-AI collaboration (HAIC) in decision-making aims to create synergistic teaming between human decision-makers and AI systems. Learning to defer (L2D) has been presented as a promising framework to determine who among humans and AI should make which decisions in order to optimize the performance and fairness of the combined system. Nevertheless, L2D entails several often unfeasible requirements, such as the availability of predictions from humans for every instance or ground-truth labels that are independent from said humans. Furthermore, neither L2D nor alternative approaches tackle fundamental issues of deploying HAIC systems in real-world settings, such as capacity management or dealing with dynamic environments. In this paper, we aim to identify and review these and other limitations, pointing to where opportunities for future research in HAIC may lie.

Motivation & Objective

  • Clarify the limitations of the L2D framework for real-world HAIC deployments.
  • Assess how capacity, selective labels, fairness, and dynamic environments impact HAIC performance.
  • Propose future research directions to expand beyond L2D toward holistic HAIC systems.

Proposed method

  • Review and synthesize the Learning to Defer (L2D) framework and its mathematical formulation.
  • Explain how L2D optimizes assignments via a deferral model and a main classifier.
  • Discuss limitations such as need for human predictions on all training instances and lack of capacity management.
  • Analyze challenges including selective labels, multiple decision-makers, robustness, and fairness.
  • Highlight dynamic environments and non-stationarity as unaddressed factors.
  • Outline potential future research avenues beyond L2D.

Experimental results

Research questions

  • RQ1What are the structural limitations of Learning to Defer for HAIC in practice?
  • RQ2How do capacity constraints, selective labeling, and multiple experts affect HAIC performance under L2D?
  • RQ3How can HAIC systems maintain fairness and robustness in dynamic environments?
  • RQ4What alternative or complementary approaches can address data-with-and-without-human-predictions and non-stationarity?

Key findings

  • L2D has significant practical limitations, including the requirement for human predictions on every training instance and no explicit capacity management.
  • Joint training in L2D can reduce robustness and hinder advisory roles where AI provides scores or explanations to humans.
  • Deferral to multiple experts increases data collection burden and may not be feasible in real teams, especially without concurrent predictions.
  • Capacity management, selective labels, fairness, and dynamic environments are not adequately addressed by L2D and require new methodologies.
  • Non-stationary environments and concept drift pose challenges L2D does not natively handle, necessitating ongoing updates and adaptive systems.
  • The paper calls for holistic HAIC research that integrates performance, fairness, capacity constraints, and adaptability to real-world settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.