[论文解读] Partial Identification of Causal Effects Using Proxy Variables
本文提出了一种无需完整条件(completeness condition)即可使用代理变量进行因果效应部分识别的方法,该条件通常用于代理因果推断中的点识别。通过利用条件独立性和依赖结构的界,作者推导出平均处理效应和处理组效应的非光滑界,并采用LogSumExp近似和自助法推断实现实际应用。
Proximal causal inference is a recently proposed framework for evaluating causal effects in the presence of unmeasured confounding. For point identification of causal effects, it leverages a pair of so-called treatment and outcome confounding proxy variables, to identify a bridge function that matches the dependence of potential outcomes or treatment variables on the hidden factors to corresponding functions of observed proxies. Unique identification of a causal effect via a bridge function crucially requires that proxies are sufficiently relevant for hidden factors, a requirement that has previously been formalized as a completeness condition. However, completeness is well-known not to be empirically testable, and although a bridge function may be well-defined, lack of completeness, sometimes manifested by availability of a single type of proxy, may severely limit prospects for identification of a bridge function and thus a causal effect; therefore, potentially restricting the application of the proximal causal framework. In this paper, we propose partial identification methods that do not require completeness and obviate the need for identification of a bridge function. That is, we establish that proxies of unobserved confounders can be leveraged to obtain bounds on the causal effect of the treatment on the outcome even if available information does not suffice to identify either a bridge function or a corresponding causal effect of interest. Our bounds are non-smooth functionals of the observed data distribution. As a consequence, in the context of inference, we initially provide a smooth approximation of our bounds. Subsequently, we leverage bootstrap confidence intervals on the approximated bounds. We further establish analogous partial identification results in related settings where identification hinges upon hidden mediators for which proxies are available.
研究动机与目标
- 解决代理因果推断的局限性,后者通常需要不可检验的完整条件才能实现因果效应的点识别。
- 在仅有一个或代理变量不足的丰富度的场景中实现因果效应估计,使点识别不可行。
- 开发在混淆桥函数因缺乏完整性而无法识别时,仍能提供有效界的方法。
- 将代理因果推断的适用性扩展到代理变量有限或代理相关性较弱的现实场景。
- 为这些界提供基于平滑近似(LogSumExp)和自助置信区间的推断程序。
提出的方法
- 基于条件独立性假设和代理变量关系,推导平均处理效应(ATE)和处理组效应(ETT)的非光滑函数界。
- 使用LogSumExp近似对非光滑界进行平滑处理,以实现可微性和数值稳定性。
- 应用自助重抽样方法构建近似界周围的置信区间,确保有效的频率学推断。
- 利用涉及处理、结果、代理变量和未观测混淆因子的条件独立结构,推导界,而无需识别桥函数。
- 基于联合密度与乘积密度之比(例如,$\frac{p(w,z|a,x)}{p(w|a,x)p(z|a,x)}$)构建界,以捕捉代理变量之间的依赖关系。
- 结合基于结果变量支持集的平凡界,确保在代理依赖关系微弱时最终界仍然有效。

实验结果
研究问题
- RQ1当代理变量的完整性条件不成立、无法实现桥函数点识别时,是否仍可对因果效应进行部分识别?
- RQ2在仅有一个代理变量或代理丰富度有限的场景中,如何推导平均处理效应的界?
- RQ3放松完整性假设对使用代理变量进行因果效应估计的可行性与有效性有何影响?
- RQ4是否可以使用平滑近似和重抽样方法构建部分识别因果效应的可靠推断?
- RQ5当代理变量较弱或不完整时,所提界在有限样本中与替代方法相比表现如何?
主要发现
- 本文证明,即使在缺乏完整性条件下,也可通过代理变量推导出平均处理效应和处理组效应的界,实现因果效应的部分识别。
- 这些界是非光滑的观测分布泛函,基于条件独立性和代理依赖结构推导得出。
- LogSumExp近似为非光滑界提供了平滑、可微的近似,从而实现实际估计与推断。
- 围绕近似界构建了自助置信区间,确保对部分识别因果参数的频率学覆盖有效性。
- 即使仅有一种代理变量可用,该方法依然有效,从而降低了对同时具备处理和结果混淆代理变量的需求。
- 理论结果表明,当代理依赖较强时,所提界比平凡支持界更紧,从而在现实场景中提升了估计精度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。