[论文解读] Psy-LLM: Scaling up Global Mental Health Psychological Services with AI-based Large Language Models
Psy-LLM 提出一种基于 AI 的前端工具,用于在线心理咨询,通过在专业问答和抓取的心理学数据上对 PanGu 和 WenZhong 进行微调,在困惑度和人工评估下进行评估,以帮助临床医生并筛查紧急病例。
The demand for psychological counselling has grown significantly in recent years, particularly with the global outbreak of COVID-19, which has heightened the need for timely and professional mental health support. Online psychological counselling has emerged as the predominant mode of providing services in response to this demand. In this study, we propose the Psy-LLM framework, an AI-based assistive tool leveraging Large Language Models (LLMs) for question-answering in psychological consultation settings to ease the demand for mental health professions. Our framework combines pre-trained LLMs with real-world professional Q\&A from psychologists and extensively crawled psychological articles. The Psy-LLM framework serves as a front-end tool for healthcare professionals, allowing them to provide immediate responses and mindfulness activities to alleviate patient stress. Additionally, it functions as a screening tool to identify urgent cases requiring further assistance. We evaluated the framework using intrinsic metrics, such as perplexity, and extrinsic evaluation metrics, with human participant assessments of response helpfulness, fluency, relevance, and logic. The results demonstrate the effectiveness of the Psy-LLM framework in generating coherent and relevant answers to psychological questions. This article discusses the potential and limitations of using large language models to enhance mental health support through AI technologies.
研究动机与目标
- 通过使 AI 辅助的在线心理咨询成为可能,缓解执业心理健康专业人员短缺。
- 开发面向专业人员的前线工具,以提供即时回应和正念活动,并筛查紧急病例。
- 将预训练的中文大语言模型与领域特定的问答数据结合起来,以提升咨询质量和可及性。
- 使用内在指标(困惑度)和外在指标(人工评估)以及真实世界部署反馈来评估框架。
提出的方法
- 以 PanGu 和 WenZhong 大规模预训练的中文语言模型作为基底。
- 在 PsyQA(22,000 个问题和 56,000 个答案)以及广泛抓取的中文心理学文章上进行微调。
- 收集并清洗数据(去重、广告、短样本、URLs、用户名,以及标点符号归一化)以创建高质量的训练语料。
- 在 2.85GB 心理学语料库上训练 PanGu 350M,并用 PsyQA 数据进行微调;使用 100,000 次迭代实现收敛。
- 纳入 WenZhong-110M 以进行额外微调,并将数据处理适配模型要求(分词、最大序列长度)。
- 通过困惑度和专业心理学评估来评估数据集质量;部署一个专门的网站用于用户反馈和迭代改进。
实验结果
研究问题
- RQ1在在线咨询情境下,Psy-LLM 如何有效地产生连贯、相关且专业的心理回应?
- RQ2在高负荷时期,Psy-LLM 是否能通过减轻工作量和协助紧急病例筛查来帮助人类咨询师?
- RQ3PanGu 与 WenZhong 作为面向心理学的中文 QA 基底模型的比较影响是什么?
- RQ4基于网页的前端与 AI 辅助回应是否提高了获取性并减轻寻求心理健康支持的耻感?
主要发现
- 在服务器部署时,该框架可以在几秒钟内向用户生成回应。
- 将 PsyQA 与抓取的中文心理学数据结合的数据管道生成用于微调 PanGu 350M 和 WenZhong-110M 的训练语料。
- 以困惑度为基础的数据质量评估辅以专业心理学家评审,指导数据清洗和模型训练。
- 该数据集大约包含 400,000 条样本,主要数据来自 Tianya(约 2GB),再加上 Zhihu 和 Yixinli 的贡献,反映了领域相关内容。
- 该方法同时针对临床医生支持(辅助工具)和在人类咨询师不可用时的面向患者的在线咨询。
- 研究讨论了将大型语言模型应用于心理健康支持的潜在收益与局限性,包括伦理和可靠性方面的考量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。