Skip to main content
QUICK REVIEW

[论文解读] DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews

Sergio Burdisso, Ernesto Reyes-Ramírez|arXiv (Cornell University)|Apr 22, 2024
Mental Health Research TopicsPsychology被引用 3
一句话总结

本文研究了在自动抑郁检测中,DAIC-WOZ数据集中因治疗师提示语而产生的非预期偏差。通过消融实验与注意力分析,作者表明模型会利用这些提示语作为判别性捷径——聚焦于特定的心理健康探问问题——通过有意利用这种偏差,模型实现了0.90的F1分数,从而削弱了在真实临床环境中的泛化能力。

ABSTRACT

Automatic depression detection from conversational data has gained significant interest in recent years. The DAIC-WOZ dataset, interviews conducted by a human-controlled virtual agent, has been widely used for this task. Recent studies have reported enhanced performance when incorporating interviewer's prompts into the model. In this work, we hypothesize that this improvement might be mainly due to a bias present in these prompts, rather than the proposed architectures and methods. Through ablation experiments and qualitative analysis, we discover that models using interviewer's prompts learn to focus on a specific region of the interviews, where questions about past experiences with mental health issues are asked, and use them as discriminative shortcuts to detect depressed participants. In contrast, models using participant responses gather evidence from across the entire interview. Finally, to highlight the magnitude of this bias, we achieve a 0.90 F1 score by intentionally exploiting it, the highest result reported to date on this dataset using only textual information. Our findings underline the need for caution when incorporating interviewers' prompts into models, as they may inadvertently learn to exploit targeted prompts, rather than learning to characterize the language and behavior that are genuinely indicative of the patient's mental health condition.

研究动机与目标

  • 调查在使用治疗师提示语进行自动抑郁检测时,报告的性能提升是否源于真正的诊断学习,还是数据集偏差。
  • 分析使用治疗师提示语训练的模型与仅使用参与者回应的模型在注意力分配上的差异。
  • 证明现有研究中较高的性能表现可能源于对提示语中局部偏差的利用,而非学习到有意义的心理健康指标。
  • 强调在真实临床人工智能系统中依赖此类偏差信号所存在的伦理与实际风险。
  • 倡导采用更稳健的评估协议,以考虑临床NLP基准中基于提示语的捷径学习现象。

提出的方法

  • 通过对比仅使用参与者回应与同时结合治疗师提示语训练的模型,开展消融实验。
  • 使用注意力可视化技术分析模型在推理过程中关注的位置,识别出对特定提示区域的集中关注。
  • 训练一个模型,通过强化对心理健康探问问题的回答来有意利用提示语偏差。
  • 在DAIC-WOZ测试集上,使用标准指标(F1分数)评估性能,重点关注文本模态。
  • 对注意力图与模型行为在不同访谈片段中的表现进行定性分析。
  • 在完整访谈序列上重复分析,比较全局与局部证据收集的差异。

实验结果

研究问题

  • RQ1在抑郁检测模型中引入治疗师提示语是否因真正的诊断信号而带来性能提升,还是因数据集偏差所致?
  • RQ2使用治疗师提示语的模型与仅使用参与者回应的模型在注意力分布上存在何种差异?
  • RQ3模型是否能仅通过利用提示语的结构与内容实现高性能表现?
  • RQ4依赖提示语的模型在多大程度上会聚焦于特定访谈片段,尤其是包含心理健康探问问题的部分?
  • RQ5基于提示语的偏差在多大程度上影响了自动抑郁检测系统在真实世界环境中的泛化能力与可靠性?

主要发现

  • 使用治疗师提示语的模型会将注意力集中于访谈的特定局部区域——主要是后半部分,即心理探问问题集中的区域。
  • 相比之下,仅基于参与者回应训练的模型则从整个对话中收集证据,表现出更广泛分布且更具上下文关联性的注意力模式。
  • 通过有意利用提示语偏差,作者在DAIC-WOZ数据集上实现了0.90的F1分数,这是迄今仅使用文本信息达到的最高结果。
  • 使用提示语带来的性能提升主要归因于提示内容中存在强烈且可被利用的偏差,而非模型架构或特征学习的改进。
  • 研究结果表明,当前高性能模型可能学习的是基于访谈者行为的判别性捷径,而非患者语言或行为中体现的抑郁特征。
  • 本研究对在使用治疗师提示语时DAIC-WOZ基准结果的有效性提出质疑,呼吁在模型评估与部署中保持谨慎。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。