Skip to main content
QUICK REVIEW

[Paper Review] When Humans and Machines Make Joint Decisions: A Non-Symmetric Bandit Model.

Sebastian Bordt, Ulrike von Luxburg|arXiv (Cornell University)|Jul 9, 2020
Reinforcement Learning in Robotics7 references4 citations
TL;DR

This paper introduces a non-symmetric bandit model to formalize human-machine decision-making, particularly in medical contexts where a doctor uses a machine assistant. It shows that exploration is fundamentally difficult without assumptions; however, under policy space independence, both agents can explore independently, resolving coordination issues in joint learning.

ABSTRACT

How can humans and machines learn to make joint decisions? This has become an important question in domains such as medicine, law and finance. We approach the question from a theoretical perspective and formalize our intuitions about human-machine decision making in a non-symmetric bandit model. In doing so, we follow the example of a doctor who is assisted by a computer program. We show that in our model, exploration is generally hard. In particular, unless one is willing to make assumptions about how human and machine interact, the machine cannot explore efficiently. We highlight one such assumption, policy space independence, which resolves the coordination problem and allows both players to explore independently. Our results shed light on the fundamental difficulties faced by the interaction of humans and machines. We also discuss practical implications for the design of algorithmic decision systems.

Motivation & Objective

  • To formalize the dynamics of joint decision-making between humans and machines in high-stakes domains like medicine, law, and finance.
  • To identify fundamental theoretical barriers to efficient exploration in human-machine collaboration.
  • To analyze how interaction assumptions affect the coordination of exploration between human and machine agents.
  • To propose conditions under which both agents can explore independently without compromising learning efficiency.
  • To inform the design of algorithmic decision systems that support effective human-machine synergy.

Proposed method

  • Formalizing human-machine interaction as a non-symmetric multi-armed bandit model, where the human and machine have asymmetric roles and information access.
  • Modeling the human as a decision-maker with limited exploration capacity and the machine as a potentially more capable but constrained learner.
  • Introducing the concept of 'policy space independence' as a structural assumption to decouple the exploration strategies of the two agents.
  • Analyzing the convergence and regret properties of the joint learning process under different assumptions about interaction and information sharing.
  • Using theoretical analysis to compare learning efficiency under symmetric vs. non-symmetric settings and with/without coordination assumptions.
  • Demonstrating that without policy space independence, the machine cannot explore efficiently due to coordination failure.

Experimental results

Research questions

  • RQ1What are the fundamental challenges to efficient exploration in human-machine decision-making systems?
  • RQ2How does the asymmetry between human and machine decision-making capabilities affect learning efficiency?
  • RQ3Under what conditions can both human and machine agents explore independently without coordination overhead?
  • RQ4What structural assumptions are necessary to resolve the coordination problem in joint exploration?
  • RQ5How does policy space independence enable more effective joint learning in non-symmetric bandit settings?

Key findings

  • Exploration is fundamentally difficult in human-machine systems unless specific assumptions are made about the interaction structure.
  • Without assumptions, the machine cannot explore efficiently due to coordination problems arising from asymmetric decision-making roles.
  • The assumption of policy space independence resolves the coordination problem, enabling both agents to explore independently.
  • Under policy space independence, both human and machine can achieve efficient learning without requiring explicit coordination during exploration.
  • The model highlights that real-world algorithmic decision systems must be designed with structural assumptions about human-machine interaction to support scalable learning.
  • The results suggest that current human-machine systems often fail to exploit the full potential of joint learning due to unmodeled coordination constraints.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.