[論文レビュー] Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging
本研究では、レントゲン画像診断のワークフローにおけるAI推論の提示順序が、人間の意思決定に与える影響を調査する。1段階ワークフロー(AIを最初に提示)と2段階ワークフロー(人間の診断後にAIを提示)を比較した結果、1段階アプローチはAIとの整合性を高め、有用性の認識を向上させ、第二意見の照会を促進したが、特にAIが誤っている場合にアーキン効果のリスクが高まることが判明した。
Details of the designs and mechanisms in support of human-AI collaboration must be considered in the real-world fielding of AI technologies. A critical aspect of interaction design for AI-assisted human decision making are policies about the display and sequencing of AI inferences within larger decision-making workflows. We have a poor understanding of the influences of making AI inferences available before versus after human review of a diagnostic task at hand. We explore the effects of providing AI assistance at the start of a diagnostic session in radiology versus after the radiologist has made a provisional decision. We conducted a user study where 19 veterinary radiologists identified radiographic findings present in patients' X-ray images, with the aid of an AI tool. We employed two workflow configurations to analyze (i) anchoring effects, (ii) human-AI team diagnostic performance and agreement, (iii) time spent and confidence in decision making, and (iv) perceived usefulness of the AI. We found that participants who are asked to register provisional responses in advance of reviewing AI inferences are less likely to agree with the AI regardless of whether the advice is accurate and, in instances of disagreement with the AI, are less likely to seek the second opinion of a colleague. These participants also reported the AI advice to be less useful. Surprisingly, requiring provisional decisions on cases in advance of the display of AI inferences did not lengthen the time participants spent on the task. The study provides generalizable and actionable insights for the deployment of clinical AI tools in human-in-the-loop systems and introduces a methodology for studying alternative designs for human-AI collaboration. We make our experimental platform available as open source to facilitate future research on the influence of alternate designs on human-AI workflows.
研究の動機と目的
- AI推論の提示タイミングが、臨床的画像診断における放射線科医の診断意思決定に与える影響を理解すること。
- ワークフローの順序が、アーキンバイアス、診断性能、時間効率、AIの有用性認識に与える影響を評価すること。
- AIの早期提示が、人間とAIの共同作業のパフォーマンスや信頼性を向上させるか、あるいは損なうかを評価すること。
- 人間-AI協働ワークフローにおける使いやすさ、信頼性、認知的負荷の間の設計的トレードオフを探ること。
- 人間が関与するシステムにおいて、実臨床現場へのAIツールの導入に役立つインサイトを提供すること。
提案手法
- Webベースの実験プラットフォームを用いて、19名の獣医放射線科医を対象に制御されたユーザースタディを実施した。
- 2つのワークフローコンfigurationを採用:1段階(レントゲン画像とAI推論を同時に表示)と2段階(AI推論の提示前に一時的な人間の診断を実施)。
- 33の画像所見に対して、バイナリのAI推論(存在/非存在)と信頼度スコアを生成するアンサンブル機械学習モデルを用いた。
- 診断意思決定、タスクに要した時間、自己評価の自信度、第二意見の照会、AIの有用性認識に関するデータを収集した。
- 放射線科医とAIの診断の整合性、アーキン効果、両ワークフローにおけるパフォーマンスを分析した。
- 今後の研究を支援するため、実験プラットフォームをオープンソースとして公開した。
実験結果
リサーチクエスチョン
- RQ1AI推論の提示タイミング(人間の診断の前後)が、放射線科医の診断意思決定に与える影響は何か?
- RQ2ワークフローの順序が、AIの推奨に向けたアーキンバイアスに与える影響は何か?
- RQ3ワークフローコンfigurationが、診断性能、評価者間信頼性、タスクに要した時間に与える影響は何か?
- RQ4異なるワークフロー条件下で、放射線科医はAI推論の有用性をどのように認識しているか?
- RQ5ワークフローデザインの影響が、第二意見の照会とAI支援診断への信頼に与える意味は何か?
主な発見
- AI推論の提示が早い1段階ワークフローでは、AIの診断が誤っていなくても、放射線科医のAI推論との整合性が2段階ワークフローに比べて顕著に高かった。
- アーキンバイアスのリスクが高かっただけでなく、特に重大でない所見においては、AIへの依存度が高まったことにより、診断性能にわずかな向上が見られた。
- 2段階ワークフローの参加者は、AIの診断と食い違う場合でも、第二意見を求める傾向が低く、AIフィードバックへの関与が薄いことが示された。
- 1段階ワークフローは放射線科医からより有用性が高いと評価され、AIの診断と自分の初期判断が食い違う場合には、同僚に相談する可能性も高かった。
- 両ワークフロー間でタスクに要した時間の差は有意ではなく、早期のAI提示が認知的負荷を増加させたり、意思決定を遅らせたりしないことが示された。
- 重大または命に関わる所見では、両ワークフローにおいて放射線科医とAIの診断の整合性が高く、高リスクのケースではアーキンバイアスの影響が緩和される可能性があることが示唆された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。