[论文解读] Who Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging
本研究探讨了在放射科工作流程中,AI推理的呈现顺序如何影响放射科医生的诊断决策。研究对比了一步流程(AI优先展示)与两步流程(人类诊断后展示AI),发现一步流程虽存在更高的锚定风险,尤其当AI出错时,但能提升与AI的一致性、增强AI的感知有用性,并促进第二诊疗意见的获取。
Details of the designs and mechanisms in support of human-AI collaboration must be considered in the real-world fielding of AI technologies. A critical aspect of interaction design for AI-assisted human decision making are policies about the display and sequencing of AI inferences within larger decision-making workflows. We have a poor understanding of the influences of making AI inferences available before versus after human review of a diagnostic task at hand. We explore the effects of providing AI assistance at the start of a diagnostic session in radiology versus after the radiologist has made a provisional decision. We conducted a user study where 19 veterinary radiologists identified radiographic findings present in patients' X-ray images, with the aid of an AI tool. We employed two workflow configurations to analyze (i) anchoring effects, (ii) human-AI team diagnostic performance and agreement, (iii) time spent and confidence in decision making, and (iv) perceived usefulness of the AI. We found that participants who are asked to register provisional responses in advance of reviewing AI inferences are less likely to agree with the AI regardless of whether the advice is accurate and, in instances of disagreement with the AI, are less likely to seek the second opinion of a colleague. These participants also reported the AI advice to be less useful. Surprisingly, requiring provisional decisions on cases in advance of the display of AI inferences did not lengthen the time participants spent on the task. The study provides generalizable and actionable insights for the deployment of clinical AI tools in human-in-the-loop systems and introduces a methodology for studying alternative designs for human-AI collaboration. We make our experimental platform available as open source to facilitate future research on the influence of alternate designs on human-AI workflows.
研究动机与目标
- 了解AI推理呈现时机如何影响放射科医生在临床影像诊断中的决策。
- 评估工作流程顺序对锚定偏倚、诊断表现、时间效率及AI感知有用性的影响。
- 评估早期接触AI是否提升或削弱人机协作的性能与信任。
- 探索人机协作工作流程中可用性、可靠性与认知负荷之间的设计权衡。
- 为在真实临床环境中部署人机协同系统提供可操作的见解。
提出的方法
- 通过基于网络的实验平台,对19名兽医放射科医生开展受控用户研究。
- 采用两种工作流程配置:一步法(X光图像与AI推理同时显示)和两步法(先进行人类初步诊断,再显示AI推理)。
- 使用集成机器学习模型为33种影像学表现生成二元AI推理结果(存在/不存在)及置信度评分。
- 收集诊断决策、任务耗时、信心评分、第二诊疗意见请求以及AI感知有用性的数据。
- 分析放射科医生与AI诊断之间的一致性、锚定效应,以及两种工作流程下的表现差异。
- 将实验平台开源,以支持未来关于人机工作流程设计的研究。
实验结果
研究问题
- RQ1AI推理呈现时机(在人类诊断前或后)如何影响放射科医生的诊断决策?
- RQ2工作流程顺序对AI建议锚定偏倚有何影响?
- RQ3工作流程配置如何影响诊断表现、阅片者间一致性及任务耗时?
- RQ4放射科医生在不同工作流程条件下如何感知AI推理的有用性?
- RQ5工作流程设计对第二诊疗意见获取及对AI辅助诊断的信任有何影响?
主要发现
- 一步工作流程中的放射科医生与AI推理的一致性显著高于两步工作流程,即使AI建议错误亦如此。
- 尽管锚定风险更高,一步工作流程因更依赖AI,尤其在非关键发现上,带来了诊断表现的微弱提升。
- 两步工作流程中的参与者,在与AI意见不一致时,更少主动寻求第二诊疗意见,表明对AI反馈的参与度降低。
- 一步工作流程被放射科医生评价为更具用处,且当AI建议与初始判断冲突时,更倾向于咨询同事。
- 两种工作流程的任务耗时无显著差异,表明早期接触AI不会增加认知负荷或延缓决策。
- 对于关键或危及生命的发现,两种工作流程中放射科医生与AI的一致性均较高,表明高风险病例可能减轻锚定效应的影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。