Skip to main content
QUICK REVIEW

[论文解读] Assessing the Usability of GutGPT: A Simulation Study of an AI Clinical Decision Support System for Gastrointestinal Bleeding Risk

Colleen Chan, Kisung You|arXiv (Cornell University)|Dec 6, 2023
Artificial Intelligence in Healthcare and Education被引用 4
一句话总结

本研究通过模拟试验评估了GutGPT——一种基于大语言模型的胃肠道出血风险临床决策支持系统——在医生和医学生中的表现。结果表明,GutGPT提升了内容掌握程度和感知使用便捷性,但其对信任与接受度的影响存在矛盾,凸显了在临床环境中优化人机交互的必要性。

ABSTRACT

Applications of large language models (LLMs) like ChatGPT have potential to enhance clinical decision support through conversational interfaces. However, challenges of human-algorithmic interaction and clinician trust are poorly understood. GutGPT, a LLM for gastrointestinal (GI) bleeding risk prediction and management guidance, was deployed in clinical simulation scenarios alongside the electronic health record (EHR) with emergency medicine physicians, internal medicine physicians, and medical students to evaluate its effect on physician acceptance and trust in AI clinical decision support systems (AI-CDSS). GutGPT provides risk predictions from a validated machine learning model and evidence-based answers by querying extracted clinical guidelines. Participants were randomized to GutGPT and an interactive dashboard, or the interactive dashboard and a search engine. Surveys and educational assessments taken before and after measured technology acceptance and content mastery. Preliminary results showed mixed effects on acceptance after using GutGPT compared to the dashboard or search engine but appeared to improve content mastery based on simulation performance. Overall, this study demonstrates LLMs like GutGPT could enhance effective AI-CDSS if implemented optimally and paired with interactive interfaces.

研究动机与目标

  • 评估临床医生对GutGPT——一种基于大语言模型的胃肠道出血风险临床决策支持系统——的信任度、接受度和可用性。
  • 将GutGPT在临床决策和知识获取方面的影响与传统交互式仪表板和互联网搜索进行对比。
  • 探究对话式人工智能界面如何影响临床医生在模拟临床情境中对人工智能CDSS的感知、使用努力预期及使用意愿。
  • 评估GutGPT在提升上消化道出血管理指南掌握程度方面的教育价值。

提出的方法

  • 开展了一项包含55名参与者(包括急诊科医生、内科医生和医学生)的模拟研究。
  • 将参与者随机分为三组:GutGPT联合电子病历仪表板组、仅交互式仪表板组,以及互联网搜索联合电子病历仪表板组。
  • 基于UTAUT模型设计模拟前后的调查问卷,以测量信任度、接受度、努力预期和使用意愿。
  • 收集并分析屏幕及视频录制数据,结合定性反馈访谈,评估界面可用性。
  • 将GutGPT与经过验证的胃肠道出血风险预测机器学习模型集成,并整合基于证据的临床指南以提供管理建议。
  • 通过模拟前后的教育评估测量内容掌握程度,以评估学习成果。
Figure 1: Flowchart depicting the study protocol, where participants complete two phases. “Interface” refers to the interactive dashboard, and “search” refers to a general internet search engine. Participants complete surveys measuring outcomes before and after each phase.
Figure 1: Flowchart depicting the study protocol, where participants complete two phases. “Interface” refers to the interactive dashboard, and “search” refers to a general internet search engine. Participants complete surveys measuring outcomes before and after each phase.

实验结果

研究问题

  • RQ1与传统仪表板或搜索引擎相比,接触GutGPT如何影响临床医生对基于人工智能的临床决策支持系统的信任度?
  • RQ2与替代信息源相比,GutGPT在上消化道出血管理内容掌握程度方面提升了多少?
  • RQ3在模拟临床环境中,GutGPT对临床医生感知到的使用便捷性(努力预期)和使用意愿有何影响?
  • RQ4临床医生如何看待基于大语言模型的系统在捕捉临床决策中社会、情感和身体因素等细微差别方面的局限性?
  • RQ5模拟研究能否有效评估基于大语言模型的CDSS在实际部署前的可用性和教育影响?

主要发现

  • 使用GutGPT的参与者在努力预期方面表现出显著提升,表明其对使用便捷性的感知优于仪表板组或搜索组。
  • GutGPT组和仪表板组在内容掌握程度方面均有提升,模拟后教育评估中正确答案的比例更高。
  • 在模拟后,GutGPT组和仪表板组的AI-CDSS信任度均有所提高,但GutGPT组的提升效果在统计上并不显著更强。
  • 调查工具表现出较高的内部一致性,大多数UTAUT指标的Cronbach’s alpha值超过0.8。
  • 两组临床医生均表达了担忧,认为人工智能系统未能充分考虑临床决策中的社会、情感和身体等细微因素。
  • 研究仍在进行中,目前已纳入55名参与者,由于样本量较小且招募未完成,未进行正式的统计检验。
Figure 2: Measurement of reliability for adapted UTAUT metrics.
Figure 2: Measurement of reliability for adapted UTAUT metrics.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。