Skip to main content
QUICK REVIEW

[论文解读] Robustness and efficiency of covariate adjusted linear instrumental variable estimators

Stijn Vansteelandt, Vanessa Didelez|arXiv (Cornell University)|Oct 6, 2015
Statistical Methods and Inference参考文献 18被引用 3
一句话总结

本文提出了一类稳健且高效的两阶段工具变量估计量,通过利用双重稳健的G估计和自适应效率最大化,在模型设定错误时仍能保持一致性。研究表明,偏差减少的双重稳健估计量在有限样本中显著优于标准两阶段最小二乘法(TSLS),尤其在暴露模型为非线性或协变量设定错误时,效率提升最高可达30%,且在各种模型设定错误下表现稳定。

ABSTRACT

Two-stage least squares (TSLS) estimators and variants thereof are widely used to infer the effect of an exposure on an outcome using instrumental variables (IVs). They belong to a wider class of two-stage IV estimators, which are based on fitting a conditional mean model for the exposure, and then using the fitted exposure values along with the covariates as predictors in a linear model for the outcome. We show that standard TSLS estimators enjoy greater robustness to model misspecification than more general two-stage estimators. However, by potentially using a wrong exposure model, e.g. when the exposure is binary, they tend to be inefficient. In view of this, we study double-robust G-estimators instead. These use working models for the exposure, IV and outcome but only require correct specification of either the IV model or the outcome model to guarantee consistent estimation of the exposure effect. As the finite sample performance of the locally efficient G-estimator can be poor, we further develop G-estimation procedures with improved efficiency and robustness properties under misspecification of some or all working models. Simulation studies and a data analysis demonstrate drastic improvements, with remarkably good performance even when one or more working models are misspecified.

研究动机与目标

  • 解决在模型设定错误下,两阶段工具变量估计量在稳健性与效率之间的权衡问题。
  • 开发在暴露模型或结果模型任一设定错误时仍保持一致的估计量,从而在稳健性上超越标准TSLS。
  • 改善局部高效G估计量在有限样本中的表现,这些估计量在工作模型设定错误时可能效率低下或存在偏差。
  • 提出自适应估计程序——经验效率最大化和偏差减少的双重稳健估计——以在多重模型设定错误下提升估计的精确度和稳定性。
  • 评估协变量调整对效率和稳健性的影响,特别是在工具变量与协变量无关的情况下。

提出的方法

  • 使用双重稳健的G估计,该方法要求在给定协变量的条件下,结果模型或工具变量模型之一正确设定,以实现暴露效应的一致估计。
  • 应用经验效率最大化,通过在工具变量模型给定协变量的条件下正确设定,来提升估计精度。
  • 提出偏差减少的双重稳健估计(BR-γ 和 BR-β),即使在工具变量模型设定错误时也能最小化偏差。
  • 采用沙漏估计量(sandwich estimators)计算稳健标准误,以支持有限样本下的推断。
  • 通过模拟研究和基于Angrist与Krueger(1991)关于教育与工资数据的真实世界分析,比较不同估计量的表现。
  • 推导并比较线性IV模型下的估计量,重点聚焦于第一阶段采用灵活模型、第二阶段采用线性模型的两阶段程序。

实验结果

研究问题

  • RQ1在暴露模型设定错误时,标准两阶段最小二乘法(TSLS)估计量的稳健性与更一般的两阶段估计量相比如何?
  • RQ2当结果模型或给定协变量的工具变量模型任一设定错误时,双重稳健G估计量是否仍能保持一致性?其与TSLS相比表现如何?
  • RQ3局部高效G估计量在有限样本中的效率和偏差特性如何?在模型设定错误时能否得到改善?
  • RQ4经验效率最大化和偏差减少策略在实践中在多大程度上提升了双重稳健估计量的表现?
  • RQ5在工具变量与协变量无关的情况下,协变量的引入在多大程度上影响了IV估计量的效率和稳健性?

主要发现

  • 标准TSLS估计量对结果模型设定错误具有稳健性,但对暴露模型设定错误不稳健,尤其当暴露变量为二值或非线性时。
  • 双重稳健G估计量在给定协变量的条件下,只要结果模型或工具变量模型之一正确设定,即可实现一致估计,其稳健性范围广于TSLS。
  • 在模型正确设定时,局部高效双重稳健G估计量在效率上优于TSLS,但在模型设定错误时,其有限样本表现较差。
  • 经验效率最大化带来了显著的效率提升——尤其在给定协变量的工具变量模型正确设定时——与TSLS相比,标准误降低了30%。
  • 偏差减少的双重稳健估计量(BR-γ 和 BR-β)表现显著改善,其标准误为0.041–0.045,远低于TSLS的0.067,且在模型设定错误下置信区间保持稳定。
  • 在Angrist与Krueger(1991)数据的分析中,双重稳健估计量得出的教育效应估计更精确,为0.09–0.10(标准误0.041–0.044),而TSLS估计为0.13(标准误0.067),且置信区间更窄、更可靠。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。