[论文解读] Causal inference taking into account unobserved confounding
本文提出了一种不确定性区间(UIs),用于平均因果效应的推断,该方法同时考虑了抽样变异性和由于未观测到的混杂因素引起的偏倚,使用偏倚参数 𝜌 来量化未观测到的混杂因素。该方法通过推导出结果回归和双重稳健估计量的偏倚表达式作为 𝜌 的函数,扩展了这些方法,表明未观测到的混杂因素引起的偏倚在样本量较大时可能主导模型误设偏倚,并通过模拟研究和一项真实世界研究表明,即使经典置信区间显示具有显著性,不确定性区间仍可能得出不确定的结论。
Causal inference with observational data can be performed under an assumption of no unobserved confounders (unconfoundedness assumption). There is, however, seldom clear subject-matter or empirical evidence for such an assumption. We therefore develop uncertainty intervals for average causal effects based on outcome regression estimators and doubly robust estimators, which provide inference taking into account both sampling variability and uncertainty due to unobserved confounders. In contrast with sampling variation, uncertainty due unobserved confounding does not decrease with increasing sample size. The intervals introduced are obtained by deriving the bias of the estimators due to unobserved confounders. We are thus also able to contrast the size of the bias due to violation of the unconfoundedness assumption, with bias due to misspecification of the models used to explain potential outcomes. This is illustrated through numerical experiments where bias due to moderate unobserved confounding dominates misspecification bias for typical situations in terms of sample size and modeling assumptions. We also study the empirical coverage of the uncertainty intervals introduced and apply the results to a study of the effect of regular food intake on health. An R-package implementing the inference proposed is available.
研究动机与目标
- 为观察性数据因果推断中的关键局限性提供解决方案:即无法验证‘无未观测到的混杂因素’这一假设。
- 开发一种实用的推断框架,不仅量化抽样变异性的不确定性,还量化潜在未观测到的混杂因素引起的不确定性。
- 对比结果回归和双重稳健估计量中,由未观测到的混杂因素引起的偏倚与模型误设偏倚的大小。
- 提供一种直观且可解释的敏感性分析工具,以区间估计替代传统的假设检验,基于合理的未观测到的混杂因素水平。
提出的方法
- 引入偏倚参数 𝜌,以量化在给定观测协变量条件下,未观测到的混杂因素与潜在结果之间的相关性。
- 推导出结果回归和双重稳健估计量的偏倚表达式,作为 𝜌 的函数。
- 通过将抽样变异性与从偏倚函数导出的边界相结合,构建不确定性区间(UIs),确保覆盖概率高于名义水平。
- 利用 UIs 进行敏感性分析,通过识别使区间仍包含零的最大 |𝜌| 值,来判断结果对未观测到的混杂因素的稳健性。
- 将该方法应用于一项关于常规饮食摄入与健康关系的真实数据集,与经典置信区间的结果进行比较。
- 提供一个名为 'ui' 的 R 包,用于实现所提出的推断方法。
实验结果
研究问题
- RQ1在观察性研究中,未观测到的混杂因素偏倚如何影响平均因果效应的估计?
- RQ2在结果回归和双重稳健估计量中,未观测到的混杂因素引起的偏倚在多大程度上超过模型误设引起的偏倚?
- RQ3包含未观测到的混杂因素的不确定性区间是否能提供比经典置信区间更可靠的推断?
- RQ4在何种最大未观测到的混杂因素水平(以 |𝜌| 衡量)下,因果效应仍可被视为非零?
- RQ5不确定性区间如何以比 p 值更直观的方式传达对未观测到的混杂因素的敏感性?
主要发现
- 在数值实验中,当 |𝜌| ≥ 0.1 时,未观测到的混杂因素引起的偏倚主导了模型误设偏倚,尤其是在倾向得分不接近零或一时。
- 在关于常规饮食摄入与健康的现实研究中,假设 |𝜌| ≤ 0.02 的不确定性区间包含了零,表明尽管经典置信区间显示具有显著性,但因果效应的正向证据仍不充分。
- 使不确定性区间仍包含零的最大 |𝜌| 约为 0.01,表明 5% 显著性结果对这一程度的未观测到的混杂因素非常敏感。
- 所提出的不确定性区间具有高于名义水平的实证覆盖概率,反映出其能够有效考虑未观测到的混杂因素。
- 该方法表明,未观测到的混杂因素引起的偏倚不会随着样本量的增加而减小,与抽样变异性不同,因此是持续且关键的不确定性来源。
- R 包 'ui' 使得该推断框架在实际研究中的应用成为可能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。