Skip to main content
QUICK REVIEW

[論文レビュー] Advancing Human-AI Complementarity: The Impact of User Expertise and Algorithmic Tuning on Joint Decision Making

Kori Inkpen, Shreya Chappidi|arXiv (Cornell University)|Aug 16, 2022
Human-Automation Interaction and Safety被引用数 4
ひとこと要約

本研究では、ユーザーの熟練度とアルゴリズムのチューニングが、動脈・静脈ラベリングタスクにおける人間とAIの協働に与える影響を調査している。AIの支援は、AIを偽陰性を低減するようにチューニングした場合、中程度のパフォーマンスを持つユーザーのパフォーマンスを最も向上させる。これは人間の強みと整合している。一方、熟練ユーザーはパーソナライズされたチューニングから最も利益を受ける。初心者ユーザーはAI支援があっても限定的な利益にとどまる。

ABSTRACT

Human-AI collaboration for decision-making strives to achieve team performance that exceeds the performance of humans or AI alone. However, many factors can impact success of Human-AI teams, including a user's domain expertise, mental models of an AI system, trust in recommendations, and more. This work examines users' interaction with three simulated algorithmic models, all with similar accuracy but different tuning on their true positive and true negative rates. Our study examined user performance in a non-trivial blood vessel labeling task where participants indicated whether a given blood vessel was flowing or stalled. Our results show that while recommendations from an AI-Assistant can aid user decision making, factors such as users' baseline performance relative to the AI and complementary tuning of AI error types significantly impact overall team performance. Novice users improved, but not to the accuracy level of the AI. Highly proficient users were generally able to discern when they should follow the AI recommendation and typically maintained or improved their performance. Mid-performers, who had a similar level of accuracy to the AI, were most variable in terms of whether the AI recommendations helped or hurt their performance. In addition, we found that users' perception of the AI's performance relative on their own also had a significant impact on whether their accuracy improved when given AI recommendations. This work provides insights on the complexity of factors related to Human-AI collaboration and provides recommendations on how to develop human-centered AI algorithms to complement users in decision-making tasks.

研究の動機と目的

  • ユーザー熟練度が共同意思決定タスクにおけるAI支援の有効性に与える影響を理解すること。
  • 特定の真正陽性率および真正陰性率を目的にAIモデルをチューニングすることの、人間-AIチームパフォーマンスに与える影響を調査すること。
  • ユーザーがAIのパフォーマンスを自らのパフォーマンスと比較してどのように評価するかが、AIの推奨に従う意思決定の正確性に与える影響を検討すること。
  • マインドセットと信頼が人間-AI協働の結果をどのように形作るかを明らかにすること。
  • 人間の強みを補完し、チームパフォーマンスを向上させるAIアシスタントの設計指針を提供すること。

提案手法

  • 150回の試行を、Stall Catchersという市民科学プラットフォームで制御実験として実施し、動脈・静脈ラベリングタスクをシミュレートした。
  • 同一の全体的正答率を示すが、真正陽性率と真正陰性率のトレードオフが異なる3つのAIモデルを用いた。
  • AIなしのベースライン正答率に基づき、初心者、中程度のパフォーマンス、熟練者という3つの熟練度レベルのユーザーのパフォーマンスを測定した。
  • 信頼、マインドセット、意思決定行動を分析するため、ユーザーのフィードバックと認識データを収集した。
  • ユーザーの意思決定とAIの推奨の整合性を分析し、正確性と一貫性への影響を評価した。
  • ユーザーのパフォーマンスレベルごとにクラスタリングを行い、各クラスタ内でのAIチューニング効果を評価した。
Figure 1. Side-by-side view of ”flowing” vs ”stalled” blood vessels from the Stall Catchers game.
Figure 1. Side-by-side view of ”flowing” vs ”stalled” blood vessels from the Stall Catchers game.

実験結果

リサーチクエスチョン

  • RQ1ユーザー熟練度のレベルが、AI推奨の意思決定正確性に与える影響は何か?
  • RQ2AIモデルの真正陽性率と真正陰性率をチューニングすることで、人間-AI協働におけるチームパフォーマンスにどのような影響を与えるか?
  • RQ3ユーザーがAIのパフォーマンスを自らのパフォーマンスと比較して評価する際、その認識がAI支援の意思決定に与える影響の程度はどの程度か?
  • RQ4人間とAIの誤りパターンの補完性が、全体のチーム正確性に与える影響は何か?
  • RQ5特定のユーザー熟練度グループのパフォーマンスを向上させるために、AIチューニングを最適化できるか?

主な発見

  • ベースライン正答率がAIとほぼ同等の、中程度のパフォーマンスを持つユーザーは、結果が最もばらつきが大きく、AI支援が効果的だったり、逆にパフォーマンスを低下させることもあった。
  • AIを偽陰性を低減するようにチューニングした場合(偽陽性が増加する代償を伴う)、ユーザーは正答率が向上した。これは、誤った推奨をより簡単に拒否できたためである。
  • 初心者ユーザーはAI支援によりパフォーマンスが向上したが、AIの正答率に達することはできず、ユーザーのベースラインパフォーマンスが低い場合、AIの恩恵は限定的であることが示された。
  • 熟練ユーザーは、AIの推奨を的確に選択的に採用することで、パフォーマンスを維持または向上させた。これは、AIの信頼性を的確に評価できる能力の高さを示している。
  • ユーザーがAIのパフォーマンスを自らのパフォーマンスと比較して評価する際、その認識がAI推奨の正確性向上に顕著に影響した。
  • チームパフォーマンスは、AIチューニングが人間の強みと整合した場合に最大に達した。特に、ユーザーが流れている血管を検出するのが得意である場合、AIは偽陰性を最小限に抑えるようにチューニングされた場合、最も効果的であった。
Figure 2. Stall Catcher interface showing feedback after a decision was made during Stage 1 and 3.
Figure 2. Stall Catcher interface showing feedback after a decision was made during Stage 1 and 3.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。