[论文解读] High-dimensional Model-assisted Inference for Local Average Treatment Effects with Instrumental Variables
本文提出了一种高维模型辅助方法,用于使用工具变量估计局部平均处理效应(LATE),通过使用Lasso惩罚的正则化校准估计来处理高维协变量中的稀疏性。该方法在较弱条件下(即工具变量倾向得分模型正确设定)确保了有效的Wald置信区间,同时允许处理和结果回归模型可能存在误设,将双重稳健推断扩展至高维设置。
Consider the problem of estimating the local average treatment effect with an instrument variable, where the instrument unconfoundedness holds after adjusting for a set of measured covariates. Several unknown functions of the covariates need to be estimated through regression models, such as instrument propensity score and treatment and outcome regression models. We develop a computationally tractable method in high-dimensional settings where the numbers of regression terms are close to or larger than the sample size. Our method exploits regularized calibrated estimation, which involves Lasso penalties but carefully chosen loss functions for estimating coefficient vectors in these regression models, and then employs a doubly robust estimator for the treatment parameter through augmented inverse probability weighting. We provide rigorous theoretical analysis to show that the resulting Wald confidence intervals are valid for the treatment parameter under suitable sparsity conditions if the instrument propensity score model is correctly specified, but the treatment and outcome regression models may be misspecified. For existing high-dimensional methods, valid confidence intervals are obtained for the treatment parameter if all three models are correctly specified. We evaluate the proposed methods via extensive simulation studies and an empirical application to estimate the returns to education.
研究动机与目标
- 开发一种在协变量数量与样本量相当或超过样本量的高维设置中,可计算处理的局部平均处理效应(LATE)估计方法。
- 通过引入带有Lasso惩罚的正则化校准估计,将双重稳健估计扩展至高维工具变量模型。
- 在较弱条件下(即工具变量倾向得分模型正确设定)建立LATE的Wald置信区间的理论有效性,同时允许处理和结果回归模型存在误设。
- 为高维模型辅助因果推断(使用工具变量)提供一个严谨的理论框架。
提出的方法
- 采用精心选择的损失函数,对三个回归模型(工具变量倾向得分、处理回归和结果回归)的系数向量进行正则化校准估计。
- 使用Lasso惩罚在稀疏性假设下实现高维设置中的稀疏估计,仅选择相关协变量。
- 以两个增广逆概率加权(AIPW)估计量之比的形式应用双重稳健估计量来估计LATE。
- 整合Tan(2020b)提出的校准估计技术,并将其适配至高维设置,以提高估计效率和稳健性。
- 在设计矩阵和误差项的正则性条件下,通过影响函数分解和浓度不等式推导估计误差的理论界。
- 在稀疏性和调参参数的适当速率条件下,建立估计量的渐近正态性。
实验结果
研究问题
- RQ1当协变量数量相对于样本量较大时,能否构建LATE的有效置信区间?
- RQ2若工具变量倾向得分模型正确设定,而处理和结果回归模型存在误设,该方法是否仍保持有效性?
- RQ3带有Lasso的正则化校准估计如何在高维工具变量模型中提升估计效率和稳健性?
- RQ4LATE的Wald置信区间在何种理论条件下具有渐近有效性?
主要发现
- 所提出的方法在工具变量倾向得分模型正确设定的条件下,即使处理和结果回归模型存在误设,也能生成有效的LATE Wald置信区间。
- 在稀疏性条件下,该方法实现了渐近正态性并保证了有效推断,即在相关回归模型中仅有少量协变量具有非零系数。
- 理论分析表明,影响函数分解中的估计误差受Lasso调参参数和稀疏性水平相关项的有界性约束,从而在适当的速率条件下确保收敛。
- 模拟研究证实,该方法在模型误设条件下仍具有稳健性和良好的覆盖性能,优于现有高维IV方法(后者要求三个模型均正确设定)。
- 对教育回报的实证应用表明,该方法在具有高维协变量的实际因果推断中具有实际应用价值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。