Skip to main content
QUICK REVIEW

[论文解读] From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models

Zachary Englhardt, Chengqian Ma|arXiv (Cornell University)|Nov 21, 2023
Digital Mental Health Interventions参考文献 57被引用 4
一句话总结

本文提出利用大语言模型(LLMs)从移动设备和行为健康数据中生成具有临床意义的见解,超越二元分类,实现人机协作的互动模式。通过在多传感器数据(如步数、睡眠时间)上应用思维链提示(chain-of-thought prompting),LLMs在抑郁症分类任务中达到61.1%的准确率,超过现有研究水平;临床医生对使用AI生成的推理与患者共同探索数据表现出强烈兴趣,有助于增强治疗联盟。

ABSTRACT

Passively collected behavioral health data from ubiquitous sensors holds significant promise to provide mental health professionals insights from patient's daily lives; however, developing analysis tools to use this data in clinical practice requires addressing challenges of generalization across devices and weak or ambiguous correlations between the measured signals and an individual's mental health. To address these challenges, we take a novel approach that leverages large language models (LLMs) to synthesize clinically useful insights from multi-sensor data. We develop chain of thought prompting methods that use LLMs to generate reasoning about how trends in data such as step count and sleep relate to conditions like depression and anxiety. We first demonstrate binary depression classification with LLMs achieving accuracies of 61.1% which exceed the state of the art. While it is not robust for clinical use, this leads us to our key finding: even more impactful and valued than classification is a new human-AI collaboration approach in which clinician experts interactively query these tools and combine their domain expertise and context about the patient with AI generated reasoning to support clinical decision-making. We find models like GPT-4 correctly reference numerical data 75% of the time, and clinician participants express strong interest in using this approach to interpret self-tracking data.

研究动机与目标

  • 为解决传统机器学习在分析被动式、多传感器行为健康数据时的局限性,此类方法常因设备间泛化能力差及信号与心理健康关联性弱而受限。
  • 探究大语言模型(LLMs)是否能从自我追踪数据中生成具有临床意义、可解释的推理,超越二元分类,迈向可操作的洞察。
  • 评估一种人机协作模式的可行性及其感知价值,即临床医生与患者共同查询LLM以在上下文中解读行为趋势。
  • 识别在临床心理卫生场景中部署LLMs所面临的挑战,包括数据所有权、模型可靠性,以及过度依赖或模拟治疗的风险。

提出的方法

  • 采用思维链提示(chain-of-thought prompting)引导LLMs推理多传感器数据(如步数、睡眠时长)趋势与抑郁、焦虑等心理健康状况之间的关系。
  • 利用LLMs对自我追踪数据执行二元抑郁症分类,实现61.1%的准确率,优于当前最先进基准。
  • 对心理健康临床医生进行定性访谈,评估其对AI生成推理的看法,以及将此类工具整合进临床工作流程的意愿。
  • 设计一种协作交互模式,使临床医生与患者均可查询LLM,临床医生结合AI洞察与临床背景及患者病史进行综合判断。
  • 通过测量LLM正确引用输入数据数值的频率,评估其数值准确性(本研究中GPT-4的准确率为75%)。
  • 探讨架构与部署相关考量,包括本地设备推理及受苹果和谷歌隐私优先方法启发的安全数据共享模式。

实验结果

研究问题

  • RQ1LLMs能否从被动式移动设备与可穿戴传感器数据中生成具有临床相关性、可解释的推理,而不仅限于简单分类?
  • RQ2在治疗背景下,临床医生如何看待AI生成洞察在解读患者自我追踪数据时的价值?
  • RQ3LLMs在多大程度上能正确引用并推理行为数据中的数值趋势?此类推理在临床环境中的可靠性如何?
  • RQ4将LLMs整合进心理卫生实践存在哪些风险与挑战,特别是过度依赖、数据所有权及模拟治疗的问题?

主要发现

  • LLMs利用移动与行为健康数据在二元抑郁症分类任务中达到61.1%的准确率,超过当前最先进水平。
  • 临床医生对使用LLMs与患者共同探索数据表现出强烈兴趣,认为这有助于增强治疗联盟。
  • GPT-4在75%的情况下正确引用了数值输入数据,表明其在推理中具备显著的事实一致性。
  • LLMs的主要临床价值不在于分类本身,而在于其生成的推理能够支持医患对话,促进对行为趋势的共同解读。
  • 临床医生担忧对LLMs的过度依赖可能导致通用化、千篇一律的解读,甚至模拟治疗,从而削弱临床判断力。
  • 本研究强调,亟需构建安全、保护隐私的数据架构,使患者与临床医生均能主动参与并贡献于AI驱动的洞察生成。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。