[论文解读] On Deep Instrumental Variables Estimate
本文通过证明在温和条件下,使用深度神经网络(DNNs)的两阶段估计量可达到半参数效率界,为深度工具变量(Deep IV)提供了理论基础。第一阶段利用DNNs以无维度的极小极大最优收敛速率估计最优工具变量,使得第二阶段即使在高维或复杂结构的工具变量下,也能实现根n渐近正态性和效率。
The endogeneity issue is fundamentally important as many empirical applications may suffer from the omission of explanatory variables, measurement error, or simultaneous causality. Recently, \cite{hllt17} propose a "Deep Instrumental Variable (IV)" framework based on deep neural networks to address endogeneity, demonstrating superior performances than existing approaches. The aim of this paper is to theoretically understand the empirical success of the Deep IV. Specifically, we consider a two-stage estimator using deep neural networks in the linear instrumental variables model. By imposing a latent structural assumption on the reduced form equation between endogenous variables and instrumental variables, the first-stage estimator can automatically capture this latent structure and converge to the optimal instruments at the minimax optimal rate, which is free of the dimension of instrumental variables and thus mitigates the curse of dimensionality. Additionally, in comparison with classical methods, due to the faster convergence rate of the first-stage estimator, the second-stage estimator has {a smaller (second order) estimation error} and requires a weaker condition on the smoothness of the optimal instruments. Given that the depth and width of the employed deep neural network are well chosen, we further show that the second-stage estimator achieves the semiparametric efficiency bound. Simulation studies on synthetic data and application to automobile market data confirm our theory.
研究动机与目标
- 为了从理论上解释Deep IV在处理高维或复杂结构工具变量时的内生性问题上的经验成功。
- 为了建立使用深度神经网络的两阶段估计量在一般组合结构假设下可达到半参数效率界的理论基础。
- 为了证明第一阶段的DNN估计量的收敛速率在不依赖工具变量维数的情况下达到极小极大最优,从而克服维度灾难。
- 为了证明第二阶段估计量的二阶估计误差更小,且对光滑性要求弱于经典方法。
- 通过模拟实验和对汽车市场数据的应用验证理论发现。
提出的方法
- 提出两阶段估计量:第一阶段使用DNN从工具变量和内生变量中估计最优工具变量;第二阶段使用普通最小二乘法(OLS)利用估计的工具变量估计线性系数。
- 对简化形式方程施加潜在的组合结构假设,以实现第一阶段DNN估计量的无维度收敛。
- 推导出DNN估计量的收敛速率,证明其为极小极大最优,仅依赖于样本量以及深度与宽度的乘积,而不依赖于工具变量的数量。
- 在第一阶段使用全连接的ReLU深度神经网络,以灵活捕捉复杂函数形式,而无需显式结构假设。
- 在深度、宽度或两者随样本量发散的条件下,建立第二阶段估计量的渐近正态性和半参数效率。
- 采用偏差-方差分解和鞅型论证方法,控制第二阶段的估计误差,利用第一阶段DNN估计量的更快收敛特性。
实验结果
研究问题
- RQ1在存在内生性的线性工具变量模型中,两阶段框架下的深度神经网络能否实现半参数效率界?
- RQ2在高维工具变量中,第一阶段使用深度学习是否能克服维度灾难?
- RQ3与经典非参数估计量相比,基于DNN的第一阶段估计量的收敛速率在多大程度上依赖于工具变量的数量?
- RQ4在何种条件下,第二阶段估计量能实现根n渐近正态性和效率?
- RQ5与经典系列或核估计量相比,所提出的方法如何减少二阶估计误差?
主要发现
- 当深度与宽度的乘积在样本量的多项式范围内时,第一阶段DNN估计量可实现不依赖于工具变量维度的极小极大最优收敛速率。
- 当深度、宽度和工具变量数量的乘积在样本量的多项式范围内时,第二阶段估计量可达到半参数效率界。
- 由于第一阶段DNN估计量收敛更快,第二阶段估计量的二阶估计误差小于经典方法,如Cheng和Kosorok(2008)所形式化的内容。
- 与经典系列估计量相比,该方法对最优工具变量的光滑性要求更弱,后者要求光滑度超过工具变量维度的一半。
- 模拟研究和对汽车市场数据的实证应用验证了理论预测,显示在有限样本下性能更优,且对高维工具变量具有鲁棒性。
- 理论框架通过证明DNN在无需显式建模的情况下捕捉复杂结构的能力,为Deep IV的经验成功提供了理论依据,从而实现高效且一致的推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。