[论文解读] Naïve regression requires weaker assumptions than factor models to adjust for multiple cause confounding
本文表明,在比去混杂因子法(deconfounder)更弱的假设下,朴素回归(naïve regression)——即直接对多个处理变量回归结果——在多重原因设定下可实现渐近无偏。作者证明,当每个混杂因子影响无限多个处理变量时,朴素回归具有一致性,而去混杂因子法变体则需要额外的不可检验假设,且在实践中往往无法超越朴素回归,甚至产生诸如‘斯坦·李导致漫威155亿美元收入’等荒谬估计。
The empirical practice of using factor models to adjust for shared, unobserved confounders, $\mathbf{Z}$, in observational settings with multiple treatments, $\mathbf{A}$, is widespread in fields including genetics, networks, medicine, and politics. Wang and Blei (2019, WB) formalizes these procedures and develops the "deconfounder," a causal inference method using factor models of $\mathbf{A}$ to estimate "substitute confounders," $\hat{\mathbf{Z}}$, then estimating treatment effects by regressing the outcome, $\mathbf{Y}$, on part of $\mathbf{A}$ while adjusting for $\hat{\mathbf{Z}}$. WB claim the deconfounder is unbiased when there are no single-cause confounders and $\hat{\mathbf{Z}}$ is "pinpointed." We clarify pinpointing requires each confounder to affect infinitely many treatments. We prove under these assumptions, a naïve semiparametric regression of $\mathbf{Y}$ on $\mathbf{A}$ is asymptotically unbiased. Deconfounder variants nesting this regression are therefore also asymptotically unbiased, but variants using $\hat{\mathbf{Z}}$ and subsets of causes require further untestable assumptions. We replicate every deconfounder analysis with available data and find it fails to consistently outperform naïve regression. In practice, the deconfounder produces implausible estimates in WB's case study to movie earnings: estimates suggest comic author Stan Lee's cameo appearances causally contributed \$15.5 billion, most of Marvel movie revenue. We conclude neither approach is a viable substitute for careful research design in real-world applications.
研究动机与目标
- 评估朴素回归与去混杂因子法在多重原因设定下调整未观测混杂因子的理论与实证性能。
- 阐明去混杂因子法无偏的条件,特别是每个混杂因子影响无限多个处理变量(强无限混杂)的要求。
- 通过真实数据与模拟实验,检验去混杂因子法在有限样本中是否能一致优于朴素回归。
- 评估两种方法在现实世界观察性研究中的实际可行性,尤其是在复杂混杂结构的情形下。
- 挑战因子模型在调整多重原因混杂中为必要或更优的假设。
提出的方法
- 作者建立了渐近理论,表明当混杂因子Z影响无限多个处理变量(强无限混杂)时,对结果Y直接回归处理变量A的朴素回归是无偏的。
- 他们推导了在线性-线性与非线性模型下朴素估计量的偏差,证明在强无限混杂假设下偏差会消失。
- 去混杂因子法被分析为两阶段过程:第一阶段,通过因子模型从处理变量A中估计替代混杂因子Ẑ;第二阶段,对Y回归Ẑ与A的子集。
- 本文在线性与非线性设定下,将朴素回归与多种去混杂因子变体(包括惩罚、白噪声、后验均值、子集去混杂因子估计器)进行比较。
- 通过模拟与对去混杂因子法在漫威电影票房案例研究的复现,评估了有限样本下的性能表现。
- 使用后验预测模型检验来评估模型拟合度,但作者表明该方法无法可靠识别去混杂因子何时有效。
实验结果
研究问题
- RQ1在多重原因混杂存在的情况下,朴素回归在何种条件下可实现渐近无偏?
- RQ2去混杂因子法为实现无偏性是否需要比朴素回归更强的假设?
- RQ3在各种模拟与真实世界设定下,去混杂因子法是否能在有限样本中一致优于朴素回归?
- RQ4为何去混杂因子法会产生荒谬的因果估计(如斯坦·李导致155亿美元漫威收入)?这对其可靠性有何启示?
- RQ5因子模型是否为调整多重原因混杂所必需?还是在更弱假设下,朴素回归等简单方法已足够?
主要发现
- 当每个未观测混杂因子影响无限多个处理变量时,朴素回归具有一致性,该条件比去混杂因子法实现无偏性所需的假设更弱。
- 去混杂因子法需要在强无限混杂之外,额外依赖不可检验的假设,尤其在使用处理变量子集或估计的替代混杂因子时更为明显。
- 在去混杂因子法自身对漫威演员的案例研究中,该方法产生了荒谬估计:斯坦·李的客串演出被估计为对票房收入造成155亿美元的因果影响。
- 使用可用数据复现所有去混杂因子分析表明,朴素回归在有限样本中始终优于或至少不逊于去混杂因子法。
- 即使在非线性设定下,朴素回归仍能有效利用参数信息,而去混杂因子变体则需依赖强烈且不可验证的关于处理效应的假设。
- 目前,朴素回归与去混杂因子法均不适合在缺乏严谨研究设计的情况下用于现实世界的因果推断,因为两者均依赖不可检验的假设,且可能产生误导性结果。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。