Skip to main content
QUICK REVIEW

[论文解读] Cluster-Robust Standard Errors for Linear Regression Models with Many Controls

Riccardo D'Adamo|arXiv (Cornell University)|Jun 19, 2018
Statistical Methods and Bayesian Inference参考文献 24被引用 4
一句话总结

本文提出了一种针对具有大量控制变量的线性回归模型的新聚类稳健方差估计量,校正了当控制变量数量与样本量成比例增长时标准聚类稳健误差的不一致性。该方法通过调整由大量控制变量引入的偏差,在高维渐近条件下确保了有效的推断,理论分析与蒙特卡洛模拟均表明其在小样本中表现更优。

ABSTRACT

It is common practice in empirical work to employ cluster-robust standard errors when using the linear regression model to estimate some structural/causal effect of interest. Researchers also often include a large set of regressors in their model specification in order to control for observed and unobserved confounders. In this paper we develop inference methods for linear regression models with many controls and clustering. We show that inference based on the usual cluster-robust standard errors by Liang and Zeger (1986) is invalid in general when the number of controls is a non-vanishing fraction of the sample size. We then propose a new clustered standard errors formula that is robust to the inclusion of many controls and allows to carry out valid inference in a variety of high-dimensional linear regression models, including fixed effects panel data models and the semiparametric partially linear model. Monte Carlo evidence supports our theoretical results and shows that our proposed variance estimator performs well in finite samples. The proposed method is also illustrated with an empirical application that re-visits Donohue III and Levitt's (2001) study of the impact of abortion on crime.

研究动机与目标

  • 解决当控制变量数量作为样本量的非消失比例增长时,传统聚类稳健标准误的不一致性问题。
  • 为具有大量控制变量和聚类结构的线性回归模型开发一种稳健推断方法,尤其适用于高维情形。
  • 在高维渐近条件下,将有效推断扩展至固定效应面板数据模型和半参数部分线性模型。
  • 校正由于大量控制变量导致的小样本中聚类稳健标准误的偏差。
  • 提供一种即使在 $ K_n/n \not\to 0 $ 时仍保持一致性和渐近有效性的方差估计量。

提出的方法

  • 提出一种新的聚类稳健方差估计量,通过调整高维线性模型中大量控制变量引入的偏差。
  • 基于投影矩阵 $ M_n = X_n(X_n'X_n)^{-1}X_n' $ 推导出修正后的方差估计量,其中 $ X_n $ 包含处理变量和控制变量。
  • 引入一种权重方案 $ \kappa_n $,以考虑组内依赖性及误差协方差矩阵的结构。
  • 使用基于 $ S_n' (M_n \otimes M_n) S_n $ 的中间项逆矩阵的一致估计量 $ \hat{\Sigma}_n(\kappa_n^{\texttt{CR}}) $。
  • 通过集合 $ \mathcal{V}_{g,n} $ 和 $ \mathcal{R}_{g,i,n} $ 重新参数化误差结构,以定义组内非零协方差块。
  • 在高维渐近条件下,将该估计量应用于具有大量控制变量的模型,包括固定效应模型和半参数部分线性模型。

实验结果

研究问题

  • RQ1当控制变量数量随样本量成比例增长时,Liang与Zeger(1986)提出的传统聚类稳健标准误估计量是否一致?
  • RQ2能否构造一种在高维渐近条件下($ K_n/n \not\to 0 $)仍保持一致性和有效性的新聚类稳健方差估计量?
  • RQ3与标准聚类稳健误差相比,所提出的估计量在有限样本中的表现如何?
  • RQ4该新估计量能否应用于具有固定效应和半参数成分的大量控制变量模型?
  • RQ5所提出的估计量实现渐近正态性和一致性的充分条件是什么?

主要发现

  • 当 $ K_n/n \not\to 0 $ 时,传统聚类稳健标准误估计量不一致,导致高维设定下推断失效。
  • 所提出的估计量 $ \hat{\Sigma}_n(\kappa_n^{\texttt{CR}}) $ 在高维渐近条件下保持一致且渐近有效,即使 $ K_n $ 的增长速度与 $ n $ 相当。
  • 蒙特卡洛模拟证据表明,新估计量在有限样本中表现良好,减少了大小扭曲并提高了置信区间的覆盖概率。
  • 该估计量在固定效应面板模型和具有大量控制变量的半参数部分线性模型中依然有效。
  • 理论分析确认 $ \mathbb{E}[\tilde{\Sigma}_n(\mathbf{I}_{L_n})|\mathcal{X}_n,\mathcal{W}_n] = \Sigma_n + o_p(1) $,确保了一致性。
  • 该方法校正了因大量控制变量导致的标准误偏差,且不依赖于非正态近似。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。