Skip to main content
QUICK REVIEW

[논문 리뷰] Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging

Riccardo Fogliato, Shreya Chappidi|arXiv (Cornell University)|2022. 05. 19.
Artificial Intelligence in Healthcare and Education인용 수 7
한 줄 요약

이 연구는 영상의학 워크플로우에서 인공지능(AI) 추론의 순서가 인간의 의사결정에 미치는 영향을 조사한다. 일단에 워크플로우(사전에 AI를 노출)와 이중에 워크플로우(인간 진단 후 AI를 노출)를 비교하여, 일단에 워크플로우가 AI와의 일치도를 높이고, AI의 유용성 인식을 향상시키며, 제2의 의견 청취를 증가시키는 것으로 나타났다. 이는 특히 AI가 잘못되었을 경우의 앵커링 위험 증가에도 불구하고 그렇다.

ABSTRACT

Details of the designs and mechanisms in support of human-AI collaboration must be considered in the real-world fielding of AI technologies. A critical aspect of interaction design for AI-assisted human decision making are policies about the display and sequencing of AI inferences within larger decision-making workflows. We have a poor understanding of the influences of making AI inferences available before versus after human review of a diagnostic task at hand. We explore the effects of providing AI assistance at the start of a diagnostic session in radiology versus after the radiologist has made a provisional decision. We conducted a user study where 19 veterinary radiologists identified radiographic findings present in patients' X-ray images, with the aid of an AI tool. We employed two workflow configurations to analyze (i) anchoring effects, (ii) human-AI team diagnostic performance and agreement, (iii) time spent and confidence in decision making, and (iv) perceived usefulness of the AI. We found that participants who are asked to register provisional responses in advance of reviewing AI inferences are less likely to agree with the AI regardless of whether the advice is accurate and, in instances of disagreement with the AI, are less likely to seek the second opinion of a colleague. These participants also reported the AI advice to be less useful. Surprisingly, requiring provisional decisions on cases in advance of the display of AI inferences did not lengthen the time participants spent on the task. The study provides generalizable and actionable insights for the deployment of clinical AI tools in human-in-the-loop systems and introduces a methodology for studying alternative designs for human-AI collaboration. We make our experimental platform available as open source to facilitate future research on the influence of alternate designs on human-AI workflows.

연구 동기 및 목표

  • 임상 영상에서 AI 추론 제시 시점이 방사선 전문의의 진단 결정에 미치는 영향를 이해하기 위해.
  • 워크플로우 순서가 앵커링 편향, 진단 성능, 시간 효율성, AI의 유용성 인식에 미치는 영향를 평가하기 위해.
  • 초기 AI 노출이 인간-AI 팀워크 성능과 신뢰도를 향상시키는지 또는 약화시키는지 평가하기 위해.
  • 인간-AI 협업 워크플로우에서 사용성, 신뢰성, 인지 부하 간의 설계적 트레이드오���을 탐색하기 위해.
  • 실제 임상 환경에서 인간이 참여하는 시스템에 AI 도구를 구현하기 위한 실질적 통찰을 제공하기 위해.

제안 방법

  • 웹 기반 실험 플랫폼을 사용하여 19명의 수의학 영상의학 전문의를 대상으로 통제된 사용자 연구를 수행하였다.
  • 두 가지 워크플로우 구성 방식을 적용: 일단에(X레이 영상과 AI 추론을 동시에 표시) 및 이중에(인간 진단을 전제로 한 후 AI 추론 제시).
  • 33개의 레이저 영상 소견에 대해 이종 기계학습 모델을 활용해 이진 AI 추론(유무) 및 신뢰도 점수를 생성하였다.
  • 진단 결정, 작업 소요 시간, 자신감 평가, 제2의 의견 요청, AI의 유용성 인식 데이터를 수집하였다.
  • 방사선 전문의와 AI의 진단 간 일치도, 앵커링 효과, 양 워크플로우 간 성능을 분석하였다.
  • 향후 인간-AI 워크플로우 설계 연구를 지원하기 위해 실험 플랫폼을 오픈소스로 공개하였다.

실험 결과

연구 질문

  • RQ1AI 추론 제시 시점(인간 진단 이전 대비 이후)이 방사선 전문의의 진단 결정에 어떤 영향를 미치는가?
  • RQ2워크플로우 순서가 AI 권고에 대한 앵커링 편향에 어떤 영향를 미치는가?
  • RQ3워크플로우 구성이 진단 성능, 평가자 간 일치도, 작업 소요 시간에 어떤 영향를 미치는가?
  • RQ4다른 워크플로우 조건에서 방사선 전문의는 AI 추론의 유용성을 어떻게 인식하는가?
  • RQ5워크플로우 설계가 제2의 의견 청취 및 AI 보조 진단에 대한 신뢰에 어떤 영향를 미치는가?

주요 결과

  • AI 권고가 잘못되었을 경우에도, 일단에 워크플로우에 참여한 방사선 전문의는 이중에 워크플로우에 비해 AI 추론과의 일치도가 유의미하게 높았다.
  • 앵커링 위험이 높음에도 불구하고, 특히 비중대한 소견에 대해 AI에 의존도가 증가함에 따라 일기적 성능 향상이 약간 발생하였다.
  • 이중에 워크플로우에 참여한 참가자들은 AI와 의견이 맞지 않을 경우 제2의 의견을 구하는 비율이 낮아, AI 피드백에 덜 참여하는 것으로 나타났다.
  • 일단에 워크플로우는 방사선 전문의에게 더 유용하다고 평가되었으며, AI 권고가 자신의 初기 평가와 다를 경우 동료와 상의할 가능성이 더 높았다.
  • 작업 소요 시간은 두 워크플로우 간 유의미한 차이가 없었으며, 이는 초기 AI 노출이 인지 부하를 증가시키거나 의사결정을 지연시키지 않는다는 것을 시사한다.
  • 중대하거나 생명을 위협하는 소견의 경우, 두 워크플로우 모두에서 방사선 전문의와 AI 간 일치도가 높았으며, 이는 고위험 사례에서는 앵커링 효과가 완화될 수 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.