Skip to main content
QUICK REVIEW

[Paper Review] Cluster-Robust Standard Errors for Linear Regression Models with Many Controls

Riccardo D'Adamo|arXiv (Cornell University)|Jun 19, 2018
Statistical Methods and Bayesian Inference24 references4 citations
TL;DR

This paper proposes a new cluster-robust variance estimator for linear regression models with many controls, correcting the inconsistency of standard cluster-robust errors when the number of controls grows proportionally with sample size. The method ensures valid inference under high-dimensional asymptotics by adjusting for the bias introduced by many control variables, with theoretical and Monte Carlo support showing improved finite-sample performance.

ABSTRACT

It is common practice in empirical work to employ cluster-robust standard errors when using the linear regression model to estimate some structural/causal effect of interest. Researchers also often include a large set of regressors in their model specification in order to control for observed and unobserved confounders. In this paper we develop inference methods for linear regression models with many controls and clustering. We show that inference based on the usual cluster-robust standard errors by Liang and Zeger (1986) is invalid in general when the number of controls is a non-vanishing fraction of the sample size. We then propose a new clustered standard errors formula that is robust to the inclusion of many controls and allows to carry out valid inference in a variety of high-dimensional linear regression models, including fixed effects panel data models and the semiparametric partially linear model. Monte Carlo evidence supports our theoretical results and shows that our proposed variance estimator performs well in finite samples. The proposed method is also illustrated with an empirical application that re-visits Donohue III and Levitt's (2001) study of the impact of abortion on crime.

Motivation & Objective

  • To address the inconsistency of conventional cluster-robust standard errors when the number of control variables grows as a non-vanishing fraction of sample size.
  • To develop a robust inference method for linear regression models with many controls and clustering, particularly in high-dimensional settings.
  • To extend valid inference to models such as fixed effects panel data and semiparametric partially linear models under high-dimensional asymptotics.
  • To correct the small-sample bias in cluster-robust standard errors that arises due to many control variables.
  • To provide a variance estimator that remains consistent and asymptotically valid even when $ K_n/n \not\to 0 $.

Proposed method

  • Proposes a new cluster-robust variance estimator that adjusts for the bias introduced by many control variables in high-dimensional linear models.
  • Derives a modified variance estimator based on the projection matrix $ M_n = X_n(X_n'X_n)^{-1}X_n' $, where $ X_n $ includes both treatment and control variables.
  • Introduces a weighting scheme $ \kappa_n $ that accounts for within-cluster dependence and the structure of the error covariance matrix.
  • Uses a consistent estimator $ \hat{\Sigma}_n(\kappa_n^{\texttt{CR}}) $ based on the inverse of the sandwich estimator's middle term involving $ S_n' (M_n \otimes M_n) S_n $.
  • Employs a reparameterization of the error structure using sets $ \mathcal{V}_{g,n} $ and $ \mathcal{R}_{g,i,n} $ to define non-zero covariance blocks within clusters.
  • Applies the estimator to models with many controls, including fixed effects and semiparametric partially linear models, under high-dimensional asymptotics.

Experimental results

Research questions

  • RQ1Is the conventional cluster-robust standard error estimator by Liang and Zeger (1986) consistent when the number of control variables grows proportionally with sample size?
  • RQ2Can a new cluster-robust variance estimator be constructed that remains consistent and valid under high-dimensional asymptotics where $ K_n/n \not\to 0 $?
  • RQ3How does the proposed estimator perform in finite samples compared to standard cluster-robust errors?
  • RQ4Can the new estimator be applied to models with fixed effects and semiparametric components under many controls?
  • RQ5What are the sufficient conditions under which the proposed estimator achieves asymptotic normality and consistency?

Key findings

  • The conventional cluster-robust standard error estimator is inconsistent when $ K_n/n \not\to 0 $, invalidating inference in high-dimensional settings.
  • The proposed estimator $ \hat{\Sigma}_n(\kappa_n^{\texttt{CR}}) $ is consistent and asymptotically valid under high-dimensional asymptotics, even when $ K_n $ grows as fast as $ n $.
  • Monte Carlo evidence shows the new estimator performs well in finite samples, reducing size distortions and improving coverage of confidence intervals.
  • The estimator remains valid in fixed effects panel models and semiparametric partially linear models with many controls.
  • Theoretical analysis confirms that $ \mathbb{E}[\tilde{\Sigma}_n(\mathbf{I}_{L_n})|\mathcal{X}_n,\mathcal{W}_n] = \Sigma_n + o_p(1) $, ensuring consistency.
  • The method corrects for bias in standard errors arising from many control variables without relying on non-Gaussian approximations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.