Skip to main content
QUICK REVIEW

[Paper Review] Advancing Human-AI Complementarity: The Impact of User Expertise and Algorithmic Tuning on Joint Decision Making

Kori Inkpen, Shreya Chappidi|arXiv (Cornell University)|Aug 16, 2022
Human-Automation Interaction and Safety4 citations
TL;DR

This study investigates how user expertise and algorithmic tuning affect human-AI collaboration in a blood vessel labeling task. It finds that AI assistance boosts performance most for mid-performing users when the AI is tuned to reduce false negatives—aligning with human strengths—while expert users benefit most from personalized tuning, and novice users see limited gains despite AI support.

ABSTRACT

Human-AI collaboration for decision-making strives to achieve team performance that exceeds the performance of humans or AI alone. However, many factors can impact success of Human-AI teams, including a user's domain expertise, mental models of an AI system, trust in recommendations, and more. This work examines users' interaction with three simulated algorithmic models, all with similar accuracy but different tuning on their true positive and true negative rates. Our study examined user performance in a non-trivial blood vessel labeling task where participants indicated whether a given blood vessel was flowing or stalled. Our results show that while recommendations from an AI-Assistant can aid user decision making, factors such as users' baseline performance relative to the AI and complementary tuning of AI error types significantly impact overall team performance. Novice users improved, but not to the accuracy level of the AI. Highly proficient users were generally able to discern when they should follow the AI recommendation and typically maintained or improved their performance. Mid-performers, who had a similar level of accuracy to the AI, were most variable in terms of whether the AI recommendations helped or hurt their performance. In addition, we found that users' perception of the AI's performance relative on their own also had a significant impact on whether their accuracy improved when given AI recommendations. This work provides insights on the complexity of factors related to Human-AI collaboration and provides recommendations on how to develop human-centered AI algorithms to complement users in decision-making tasks.

Motivation & Objective

  • To understand how user expertise influences the effectiveness of AI assistance in joint decision-making tasks.
  • To investigate how tuning AI models for specific true positive and true negative rates affects human-AI team performance.
  • To examine how users' perceptions of AI performance relative to their own impact decision accuracy when using AI recommendations.
  • To explore the role of mental models and trust in shaping human-AI collaboration outcomes.
  • To provide design guidelines for creating AI assistants that complement human strengths and improve team performance.

Proposed method

  • Conducted a controlled experiment with 150 trials on the Stall Catchers citizen science platform, simulating a blood vessel labeling task.
  • Used three AI models with identical overall accuracy but different trade-offs between true positive and true negative rates.
  • Measured user performance across three expertise levels: novice, mid-performer, and expert, based on baseline accuracy without AI.
  • Collected user feedback and perception data to analyze trust, mental models, and decision-making behavior.
  • Analyzed alignment between user decisions and AI recommendations to assess impact on accuracy and consistency.
  • Used clustering to group users by performance level and evaluated AI tuning effects within each cluster.
Figure 1. Side-by-side view of ”flowing” vs ”stalled” blood vessels from the Stall Catchers game.
Figure 1. Side-by-side view of ”flowing” vs ”stalled” blood vessels from the Stall Catchers game.

Experimental results

Research questions

  • RQ1How does user expertise level affect the impact of AI recommendations on decision accuracy?
  • RQ2How does tuning an AI model’s true positive vs. true negative rate influence team performance in human-AI collaboration?
  • RQ3To what extent do users’ perceptions of AI performance relative to their own affect their decision-making when using AI assistance?
  • RQ4How does the complementarity between human and AI error patterns affect overall team accuracy?
  • RQ5Can AI tuning be optimized to enhance performance for specific user expertise groups?

Key findings

  • Mid-performing users, whose baseline accuracy was similar to the AI’s, showed the most variable outcomes—AI assistance helped in some cases but hurt performance in others.
  • When the AI was tuned to reduce false negatives (at the cost of more false positives), users improved in accuracy because they could more easily reject incorrect recommendations.
  • Novice users improved with AI assistance but did not reach the AI’s accuracy level, indicating limited benefit from AI when user baseline performance is low.
  • Highly proficient users maintained or improved their performance by selectively following AI recommendations, suggesting strong ability to assess AI reliability.
  • Users’ perception of the AI’s performance relative to their own significantly influenced whether AI recommendations improved their accuracy.
  • Team performance was maximized when AI tuning aligned with human strengths—particularly when users were better at detecting flowing vessels, and the AI was tuned to minimize false negatives.
Figure 2. Stall Catcher interface showing feedback after a decision was made during Stage 1 and 3.
Figure 2. Stall Catcher interface showing feedback after a decision was made during Stage 1 and 3.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.