Skip to main content
QUICK REVIEW

[论文解读] Recommendations for Bayesian hierarchical model specifications for case-control studies in mental health

Vincent Valton, Toby Wise|arXiv (Cornell University)|Nov 3, 2020
Mental Health Research Topics参考文献 5被引用 13
一句话总结

本研究建议在计算精神病学任务中,对病例对照组分别拟合贝叶斯分层模型,表明与使用共享先验合并组别相比,分别设置组级别先验能更准确、更稳健且无偏地恢复真实效应量,尤其是在数据质量较差的情况下。后者系统性地低估效应量并增加假阴性率。

ABSTRACT

Hierarchical model fitting has become commonplace for case-control studies of cognition and behaviour in mental health. However, these techniques require us to formalise assumptions about the data-generating process at the group level, which may not be known. Specifically, researchers typically must choose whether to assume all subjects are drawn from a common population, or to model them as deriving from separate populations. These assumptions have profound implications for computational psychiatry, as they affect the resulting inference (latent parameter recovery) and may conflate or mask true group-level differences. To test these assumptions we ran systematic simulations on synthetic multi-group behavioural data from a commonly used multi-armed bandit task (reinforcement learning task). We then examined recovery of group differences in latent parameter space under the two commonly used generative modelling assumptions: (1) modelling groups under a common shared group-level prior (assuming all participants are generated from a common distribution, and are likely to share common characteristics); (2) modelling separate groups based on symptomatology or diagnostic labels, resulting in separate group-level priors. We evaluated the robustness of these approaches to variations in data quality and prior specifications on a variety of metrics. We found that fitting groups separately (assumptions 2), provided the most accurate and robust inference across all conditions. Our results suggest that when dealing with data from multiple clinical groups, researchers should analyse patient and control groups separately as it provides the most accurate and robust recovery of the parameters of interest.

研究动机与目标

  • 评估不同分层模型设定对病例对照精神卫生研究中组间差异恢复准确性的影响。
  • 解决在患者组与对照组使用共享先验与独立先验时,假阳性与假阴性结果之间的权衡问题。
  • 量化数据质量(受试者数量与试验次数)在不同建模假设下对模型性能的影响。
  • 为计算精神病学中的模型设定提供基于证据的最佳实践建议。

提出的方法

  • 从奖励概率变化的多臂老虎机强化学习任务中生成模拟数据,以确保潜在参数的稳定恢复。
  • 生成了36个合成数据集,其病例组与对照组在潜在参数空间(学习率)中的重叠程度各不相同。
  • 使用贝塔先验将学习率约束在0到1之间,并通过调整聚集参数来模拟不同水平的组间相似性。
  • 使用Stan拟合模型,设置2条马尔可夫链,每条链采样3000次(其中1000次为预烧期),确保所有模型中组级别先验的校准一致。
  • 通过F1分数、假阳性/假阴性率以及效应量恢复的绝对误差(Cohen’s d)评估模型性能。
  • 通过减少受试者数量(50 → 15)和试验次数(200 → 40)对数据进行扰动,以测试在数据质量较差条件下的鲁棒性。

实验结果

研究问题

  • RQ1在共同的组级别先验下对病例组与对照组进行建模,是否比使用独立先验能更准确地恢复真实组间差异?
  • RQ2在共享与独立组级别先验模型之间,假阳性和假阴性率有何不同?
  • RQ3在受试者或试验数量较少的情况下,数据质量如何影响每种建模方法的效应量恢复准确性?
  • RQ4哪种模型设定在不同数据质量条件下能提供最稳健且无偏的效应量估计?

主要发现

  • 模型2(独立组级别先验)在恢复真实组间差异方面的F1分数达到98.26%,高于模型1的96.73%,表明整体准确性更优。
  • 模型1的假阴性率显著更高(6.03%),而模型2仅为1.75%,表明其低估了真实组间差异。
  • 在低试验条件下,模型1对真实效应量的低估最高可达-64%;而模型2在最坏情况下仅出现+29.09%的误差。
  • 模型2在所有数据质量组合下均表现出更强的鲁棒性,效应量恢复的绝对误差始终更低。
  • 模型1通过共享先验施加的正则化导致效应量系统性低估,尤其在数据质量较差时,增加了遗漏真实效应的风险。
  • 尽管模型2的假阳性率略高(2.66% vs. 0.48%),但其更高的敏感性和准确性使其成为精神卫生研究的首选。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。