Skip to main content
QUICK REVIEW

[论文解读] LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation

Siyuan Chen, Mengyue Wu|arXiv (Cornell University)|May 23, 2023
Digital Mental Health Interventions被引用 38
一句话总结

这篇论文评估以 ChatGPT 为驱动的医生与患者对话机器人在精神病诊断对话中的表现,采用迭代式提示设计并结合精神科医生与患者的人类+自动评估。

ABSTRACT

Empowering chatbots in the field of mental health is receiving increasing amount of attention, while there still lacks exploration in developing and evaluating chatbots in psychiatric outpatient scenarios. In this work, we focus on exploring the potential of ChatGPT in powering chatbots for psychiatrist and patient simulation. We collaborate with psychiatrists to identify objectives and iteratively develop the dialogue system to closely align with real-world scenarios. In the evaluation experiments, we recruit real psychiatrists and patients to engage in diagnostic conversations with the chatbots, collecting their ratings for assessment. Our findings demonstrate the feasibility of using ChatGPT-powered chatbots in psychiatric scenarios and explore the impact of prompt designs on chatbot behavior and user experience.

研究动机与目标

  • 将任务正式定义为精神科门诊诊断中的医生与患者对话机器人。
  • 与精神科医生共同设计提示,使机器人行为与现实诊断保持一致。
  • 开发并应用一个以人为本的评估框架,结合用户研究和自动评估指标。
  • 展示提示设计如何影响聊天机器人同理心、深入提问和用户体验。

提出的方法

  • 在合作精神科医生的引导下进行迭代提示设计,以塑造医生与患者对话机器人。
  • 三阶段开发:目标识别(Phase 1)、提示设计与评估框架(Phase 2)、以及与精神科医生和患者的真实用户评估(Phase 3)。
  • 对多种提示变体的实证比较(医生的 D1–D4;患者的 P1–P2)。
  • 将人工评估指标(医生的流畅性、同理心、专业能力、参与度;患者的相似性、合理性)与自动评估指标(诊断准确性、症状回忆、深入度、Distinct-1 等)相结合。
  • 在医生提示中融入同理心和深入提问;患者提示中使用明确的诚实性以及口语化/生活情境语言。
Figure 1: The overview of the psychiatrist-guided three-phase study.
Figure 1: The overview of the psychiatrist-guided three-phase study.

实验结果

研究问题

  • RQ1能否让 ChatGPT 驱动的医生与患者对话机器人逼近真实的精神科诊断对话?
  • RQ2不同的提示设计如何影响聊天机器人同理心、深入提问和用户体验?
  • RQ3人类临床医生式行为与自动评估指标在聊天机器人诊断任务中的关系是什么?
  • RQ4在表达的症状和互动风格方面,模拟患者与真实患者有何差异?
  • RQ5哪种评估框架最能体现精神科诊断对话的质量?

主要发现

  • 含同理心的医生提示能提升同理心分数,但过于重复的同理心可能影响用户体验。
  • 本研究中,未包含明确症状维度提示(D3)的提示设计在医生对话机器人中达到最高诊断准确性(55.56%),其他指标随提示而异。
  • 提示变体影响提问深度和症状回忆,D3 展现更深层次的深入提问和更高的症状准确性,但症状回忆较低。
  • 纳入抗辩与口语化语言(P2)的患者提示获得更高的现实感评分和精神科医生更佳的表达风格,尽管有时以未提及症状比例为代价。
  • 人类医生在症状覆盖方面比聊天机器人更平衡、全面,凸显在多病筛查与筛查策略方面的仍存差距。
  • 自动评估指标揭示了语言风格(Distinct-1)与患者症状报告准确性之间的权衡。
Figure 2: The iterative development process of the prompt of doctor chatbots. Psychiatrists will identify the limitations of the current version, and we will address these issues in the subsequent version.
Figure 2: The iterative development process of the prompt of doctor chatbots. Psychiatrists will identify the limitations of the current version, and we will address these issues in the subsequent version.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。