[论文解读] Large Language Models in Mental Health Care: a Scoping Review
本次范围性综述分析了在心理健康护理领域的34项关于大型语言模型的研究,旨在绘制应用、数据集、训练方法、伦理与验证差距的地图。
Objectieve:This review aims to deliver a comprehensive analysis of Large Language Models (LLMs) utilization in mental health care, evaluating their effectiveness, identifying challenges, and exploring their potential for future application. Materials and Methods: A systematic search was performed across multiple databases including PubMed, Web of Science, Google Scholar, arXiv, medRxiv, and PsyArXiv in November 2023. The review includes all types of original research, regardless of peer-review status, published or disseminated between October 1, 2019, and December 2, 2023. Studies were included without language restrictions if they employed LLMs developed after T5 and directly investigated research questions within mental health care settings. Results: Out of an initial 313 articles, 34 were selected based on their relevance to LLMs applications in mental health care and the rigor of their reported outcomes. The review identified various LLMs applications in mental health care, including diagnostics, therapy, and enhancing patient engagement. Key challenges highlighted were related to data availability and reliability, the nuanced handling of mental states, and effective evaluation methods. While LLMs showed promise in improving accuracy and accessibility, significant gaps in clinical applicability and ethical considerations were noted. Conclusion: LLMs hold substantial promise for enhancing mental health care. For their full potential to be realized, emphasis must be placed on developing robust datasets, development and evaluation frameworks, ethical guidelines, and interdisciplinary collaborations to address current limitations.
研究动机与目标
- 调查数据集类型、模型、训练技术及其对心理健康任务的适用性。
- 描述由LLMs支持的心理健康应用(诊断、治疗、参与、筛查、教育)。
- 确定验证措施、性能指标和评估实践。
- 考察在心理健康护理中部署LLMs的伦理、隐私、安全和监管挑战。
- 突出当前工具与临床可行性之间的差距,以指导未来工作。
提出的方法
- 遵循 PRISMA 2020 指南进行范围综述。
- 于 2023 年 11 月在 PubMed、Web of Science、Google Scholar、arXiv、medRxiv、PsyArXiv 进行全面检索。
- 初步识别出 313 篇文献;筛选后有 34 篇符合纳入标准。
- GPT-4 在题名/摘要筛选中担任二审,Cohen’s Kappa 约为 0.90,与人类评审者相比。
- 将出版物分为数据集/基准、模型开发/微调、应用/评估以及伦理/安全。
- 区分基于提示的和微调的LLMs;强调指令微调(IFT)和提示调优策略。

实验结果
研究问题
- RQ1在心理健康任务中使用的有哪些数据集和模型?
- RQ2LLMs 处理了哪些心理健康应用,以及它们如何被验证?
- RQ3在心理健康护理中,LLMs 的伦理、隐私、安全和治理考虑有哪些?
- RQ4当前 LLM 工具与临床可行性之间存在哪些差距,需要怎样的措施来弥合?
主要发现
- LLMs 被应用于对话代理、富同理心的对话、筛查,以及为患者和临床医生提供支持的工具。
- 大多数研究使用 2022–2023 年的出版物,提示调优和面向应用的工作激增;数据集/基准论文较少。
- 数据集主要来自社交媒体,亦有部分临床医生生成的对话和合成数据;许可通常为非商业用途。
- 评估在很大程度上依赖于自动化指标,如 F1、准确度、召回率和精确度,缺乏标准化的临床验证。
- 伦理、隐私和安全关注尚未得到充分研究,表明需要健全的数据治理和跨学科协作。
- 总体而言,LLMs 在诊断和患者支持方面显示出潜力,但临床可行性和伦理整合需要进一步发展。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。