[论文解读] DepreSym: A Depression Symptom Annotated Corpus and the Role of LLMs as Assessors of Psychological Markers
本论文提出 DepreSym,一个包含 21,580 个句子的语料库,这些句子针对 21 种贝克抑郁量表-II(BDI-II)症状进行了相关性标注,数据源自 37 个搜索系统所采用的聚合策略及专业人类评估员的协作。该研究评估了 GPT-4 和 ChatGPT 作为标注者的潜力,发现 GPT-4 达到了中等到高的一致性(kappa = 0.53),并通过筛选非相关句子,可将人工标注工作量减少约 68%,从而实现高效的人机协同标注流程,适用于抑郁症状检测任务。
Computational methods for depression detection aim to mine traces of depression from online publications posted by Internet users. However, solutions trained on existing collections exhibit limited generalisation and interpretability. To tackle these issues, recent studies have shown that identifying depressive symptoms can lead to more robust models. The eRisk initiative fosters research on this area and has recently proposed a new ranking task focused on developing search methods to find sentences related to depressive symptoms. This search challenge relies on the symptoms specified by the Beck Depression Inventory-II (BDI-II), a questionnaire widely used in clinical practice. Based on the participant systems' results, we present the DepreSym dataset, consisting of 21580 sentences annotated according to their relevance to the 21 BDI-II symptoms. The labelled sentences come from a pool of diverse ranking methods, and the final dataset serves as a valuable resource for advancing the development of models that incorporate depressive markers such as clinical symptoms. Due to the complex nature of this relevance annotation, we designed a robust assessment methodology carried out by three expert assessors (including an expert psychologist). Additionally, we explore here the feasibility of employing recent Large Language Models (ChatGPT and GPT4) as potential assessors in this complex task. We undertake a comprehensive examination of their performance, determine their main limitations and analyze their role as a complement or replacement for human annotators.
研究动机与目标
- 为解决社交媒体中抑郁症状检测缺乏细粒度、症状级别的标注数据的问题。
- 通过聚焦于临床相关的抑郁症状而非一般语言特征,提升模型的可解释性与泛化能力。
- 评估大型语言模型(LLMs)如 GPT-4 和 ChatGPT 在复杂相关性标注任务中作为评估者的可行性。
- 提出一种稳健的多标注者评估方法,结合临床心理学家,以确保高质量的基准数据。
- 提出一种混合标注策略,利用 LLM 对非相关句子进行预筛选,从而在保持数据质量的同时显著降低人工工作负担。
提出的方法
- 从参与 CLEF 2023 eRisk 实验任务的 37 个独立排名系统中,构建候选句子的聚合集合。
- 由三位专家评估员(包括一名持证临床心理学家)根据明确的症状特定内容,对句子是否相关进行标注。
- 将相关性定义为话题一致性以及对 BDI-II 症状中个体状态的明确提及。
- 应用正式的标注指南,并使用 Cohen’s kappa 和 Krippendorff’s alpha 进行评分者间一致性分析。
- 通过将 LLM 的预测结果与人工共识进行对比,评估 GPT-4 和 ChatGPT 作为自动化评估工具的表现。
- 提出一种混合标注策略:由 LLM 对非相关句子进行预筛选,使人工标注者的工作量减少约 68%。
实验结果
研究问题
- RQ1大型语言模型(如 GPT-4 和 ChatGPT)在标注抑郁症状相关性句子时,能否与人类专家达成足够的一致性?
- RQ2不同临床专业背景的人类评估员之间的评分者间一致性如何变化?
- RQ3LLM 在不损害最终数据集质量的前提下,能在多大程度上减少人工标注工作量?
- RQ4基于单个标注者判断的系统排名,与官方基于共识的排名相比有何差异?
- RQ5使用 LLM 作为过滤器的混合标注流程,能否提升创建临床症状标注数据集的效率与可扩展性?
主要发现
- GPT-4 与人类评估员的中位 Cohen’s kappa 达到 0.53,表明其一致性为中等到良好;而 ChatGPT 的得分仅为 0.31,表明其可靠性较低。
- 心理学家评估员的评分者间一致性最高(kappa = 0.54),且与官方共识排名的相关性最强(Kendall’s τ = 0.98)。
- GPT-4 与官方系统排名的相关性极高(τ = 0.86,τ_ap = 0.81),优于单个个体人类评估员,接近共识水平的一致性。
- 使用 GPT-4 作为预筛选工具,可将人工标注工作量减少约 68%,在 21,580 个句子的数据集中,每位人工标注者可节省约 49 小时。
- 尽管 LLM 表现优异,但仍存在显著的误报率,表明其尚不能完全替代人类专家,但作为初步筛查工具非常有效。
- 研究结果支持一种混合标注策略:即由 LLM 筛选出非相关句子,使人类专家仅聚焦于高潜力候选样本,从而提升效率与可扩展性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。