Skip to main content
QUICK REVIEW

[论文解读] Individual-level Anxiety Detection and Prediction from Longitudinal YouTube and Google Search Engagement Logs

Anis Zaman, Boyu Zhang|arXiv (Cornell University)|Jul 1, 2020
Digital Mental Health Interventions参考文献 77被引用 6
一句话总结

该论文提出了一种保护隐私且可扩展的框架,利用匿名化的纵向YouTube和Google搜索参与日志,在个体层面检测焦虑障碍并预测焦虑严重程度。通过从在线活动提取可解释的行为特征,该模型在焦虑检测任务中取得了0.83±0.09的F1分数,在预测GAD-7评分时均方误差为1.87±0.15,提供了一种低成本、非侵入性的临床监测工具。

ABSTRACT

Anxiety disorder is one of the world's most prevalent mental health conditions, arising from complex interactions of biological and environmental factors and severely interfering one's ability to lead normal life activities. Current methods for detecting anxiety heavily rely on in-person interviews, which can be expensive, time-consuming, and blocked by social stigmas. In this work, we propose an alternative method to identify individuals with anxiety and further estimate their levels of anxiety using personal online activity histories from YouTube and the Google Search engine, platforms that are used by millions of people daily. We ran a longitudinal study and collected multiple rounds of anonymized YouTube and Google Search logs from volunteering participants, along with their clinically validated ground-truth anxiety assessment scores. We then developed explainable features that capture both the temporal and contextual aspects of online behaviors. Using those, we were able to train models that (i) identify individuals having anxiety disorder with an average F1 score of 0.83 and (ii) assess the level of anxiety by predicting the gold standard Generalized Anxiety Disorder 7-item scores (ranges from 0 to 21) with a mean square error of 1.87 based on the ubiquitous individual-level online engagement data. Our proposed anxiety assessment framework is cost-effective, time-saving, scalable, and opens the door for it to be deployed in real-world clinical settings, empowering care providers and therapists to learn about anxiety disorders of patients non-invasively at any moment in time.

研究动机与目标

  • 通过减少对昂贵且带有污名化的面对面评估的依赖,应对焦虑障碍的高患病率和漏诊问题,尤其是在大学生群体中。
  • 克服基于公开社交媒体的心理健康检测方法的局限性,后者存在自我审查、隐私性低以及对非活跃用户覆盖有限的问题。
  • 利用YouTube和Google搜索中普遍存在的、私密的在线参与日志作为心理状态的代理指标,实现实时、非侵入性的持续监测。
  • 开发一种可扩展、可解释且具有临床相关性的框架,支持在真实医疗环境中实现焦虑的早期检测和长期追踪。

提出的方法

  • 开展了一项纵向研究,包含两次数据收集阶段,间隔五个月,参与者为分享了匿名YouTube和Google搜索日志的大学生。
  • 收集了经临床验证的GAD-7评分(0–21)作为焦虑严重程度的真实标签,确保与既定精神病学标准一致。
  • 设计了可解释的、低维的特征向量,捕捉时间模式(如搜索频率、会话时长)和上下文线索(如主题聚类、查询的情感极性)等在线行为特征。
  • 基于提取的特征,训练了监督机器学习模型——具体包括二分类任务(焦虑 vs. 无焦虑)和回归任务(GAD-7评分预测)。
  • 采用特征重要性分析以确保可解释性,使临床医生能够理解哪些行为模式对预测结果有贡献。
  • 实施了自愿参与、经机构审查委员会(IRB)批准的数据收集协议,确保参与者拥有完全控制权,包括数据删除权利,以保障伦理合规性和隐私保护。

实验结果

研究问题

  • RQ1仅基于YouTube和Google搜索的个体在线参与日志,能否可靠地在个体层面上预测焦虑障碍?
  • RQ2搜索和观看行为中的时间模式与上下文模式在多大程度上可作为焦虑严重程度的代理指标?
  • RQ3与临床评估相比,基于匿名化、私密在线日志训练的模型在预测金标准GAD-7评分时的准确性如何?
  • RQ4哪些关键行为特征与焦虑状态最强相关?这些特征能否被转化为临床可用的可解释形式?
  • RQ5此类系统在真实临床环境中部署的伦理与实际影响是什么?

主要发现

  • 该模型仅使用YouTube和Google搜索日志,在识别焦虑障碍个体方面实现了平均F1分数为0.83±0.09。
  • 在预测连续的GAD-7评分(0–21)时,模型实现了1.87±0.15的均方误差,表明回归性能出色。
  • 该框架展现出高度的可扩展性和成本效益,支持在临床随访之间实现远程、非侵入性的焦虑水平监测。
  • 可解释的特征——如与健康相关术语的搜索频率增加、在焦虑相关主题内容上观看YouTube的时间延长——与焦虑水平升高密切相关。
  • 纵向分析显示,两次数据收集阶段之间行为模式发生了显著变化,支持模型追踪动态焦虑状态的能力。
  • 该系统可集成至临床工作流程中,用于识别出现焦虑水平上升的患者,从而实现及时的治疗师干预和个性化护理规划。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。