[论文解读] A note on 'Collider bias undermines our understanding of COVID-19 disease risk and severity' and how causal Bayesian networks both expose and resolve the problem
本文通过使用因果贝叶斯网络建模选择偏差(特别是医护人员过度代表和吸烟者代表不足)如何扭曲观察到的关联,扩展了Griffith等人对新冠风险研究中碰撞器偏差的分析。研究表明,若缺乏适当的因果建模,研究可能错误地得出吸烟或压力会降低新冠风险的结论,而实际上,碰撞器偏差和医护人员身份的混杂导致了虚假的负相关关系。
An important recent preprint by Griffith et al highlights how 'collider bias' in studies of COVID19 undermines our understanding of the disease risk and severity. This is typically caused by the data being restricted to people who have undergone COVID19 testing, among whom healthcare workers are overrepresented. For example, collider bias caused by smokers being underrepresented in the dataset may (at least partly) explain empirical results that suggest smoking reduces the risk of COVID19. We extend the work of Griffith et al making more explicit use of graphical causal models to interpret observed data. We show that their smoking example can be clarified and improved using Bayesian network models with realistic data and assumptions. We show that there is an even more fundamental problem for risk factors like 'stress' which, unlike smoking, is more rather than less prevalent among healthcare workers; in this case, because of a combination of collider bias from the biased dataset and the fact that 'healthcare worker' is a confounding variable, it is likely that studies will wrongly conclude that stress reduces rather than increases the risk of COVID19. Indeed, "being in close contact with COVID19 people" reduces the risk of COVID19. To avoid such potentially erroneous conclusions, any analysis of observational data must take account of the underlying causal structure including colliders and confounders. If analysts fail to do this explicitly then any conclusions they make about the effect of specific risk factors on COVID19 are likely to be flawed.
研究动机与目标
- 解决基于观察性研究的新冠风险因素中碰撞器偏差的问题,特别是在仅限于检测个体的数据集中。
- 澄清并扩展Griffith等人关于选择偏差(尤其是医护人员过度代表)如何扭曲风险估计的发现。
- 证明医护人员身份的混杂可能导致风险因素(如压力)与疾病严重程度之间的真实关系发生反转。
- 展示因果贝叶斯网络如何通过显式建模潜在因果结构来揭示并解决这些偏差。
- 提醒研究人员在未考虑碰撞器偏差和混杂效应的情况下,避免从观察性数据中得出错误结论。
提出的方法
- 构建因果贝叶斯网络模型,以表示暴露变量(如吸烟、压力)、医护人员身份、检测状态与新冠结果之间的关系。
- 使用现实数据和假设,模拟当数据被限制在检测个体时,碰撞器偏差如何产生,特别是针对医疗角色的个体。
- 通过网络结构隐式应用do-演算和d-分离原理,识别并阻断由混杂引起的后门路径。
- 模拟反事实情景,评估当未测量的混杂因素(如医护人员身份)被适当地纳入考虑时,观察到的关联如何变化。
- 使用有向无环图(DAGs)可视化因果结构,说明碰撞器和混杂如何扭曲观察到的关联。
- 将有偏数据集中的观察关联与真实因果效应进行比较,以展示扭曲的程度和方向。
实验结果
研究问题
- RQ1在基于检测的数据集中,碰撞器偏差如何扭曲风险因素与新冠严重程度之间观察到的关联?
- RQ2为何研究可能错误地得出吸烟会降低新冠风险的结论,尽管生物学合理性表明并非如此?
- RQ3为何在检测数据集中医护人员的过度代表会导致与压力等风险因素之间出现虚假的负相关关系?
- RQ4因果贝叶斯网络在在多大程度上能够揭示并纠正由选择碰撞器和混杂引起的偏差?
- RQ5在分析疾病风险的观察性数据时,若未能建模潜在因果结构,会产生何种后果?
主要发现
- 在基于检测的数据集中,由于检测个体中吸烟者代表性不足,碰撞器偏差可能造成吸烟会降低新冠风险的错觉。
- 医护人员的过度代表(尤其是高压力水平的医护人员)可能导致错误地认为压力会降低而非增加疾病风险。
- 本研究表明,“与新冠患者密切接触”可能因碰撞器偏差而看似降低风险,而非基于生物学现实。
- 因果贝叶斯网络成功揭示了碰撞器偏差和混杂的存在,显示观察到的关联系统性地被反转。
- 若未显式建模因果结构,观察性研究极有可能对风险因素产生误导性结论。
- 本研究表明,即使数据准确,仅限于检测个体的抽样选择偏差也可能产生虚假的负相关关系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。