Skip to main content
QUICK REVIEW

[Paper Review] Assessing Omitted Variable Bias when the Controls are Endogenous

Paul Diegert, Matthew A. Masten|arXiv (Cornell University)|Jun 6, 2022
Culture, Economy, and Development Studies40 citations
TL;DR

The paper shows residualization-based sensitivity analyses can misrepresent robustness when controls are endogenous, and proposes new sensitivity parameters with an implementable identification framework and a Stata module.

ABSTRACT

Omitted variables are one of the most important threats to the identification of causal effects. Several widely used methods assess the impact of omitted variables on empirical conclusions by comparing measures of selection on observables with measures of selection on unobservables. The recent literature has discussed various limitations of these existing methods, however. This includes challenges that arise when the omitted variables are endogenous, meaning that they are correlated with the included controls. We develop a new approach to regression sensitivity analysis that avoids those limitations, while still allowing researchers to calibrate sensitivity parameters by comparing the magnitude of selection on observables with the magnitude of selection on unobservables as in previous methods. We illustrate our results in an empirical study of the effect of historical American frontier life on modern cultural beliefs. Finally, we implement these methods in the companion Stata module regsensitivity for easy use in practice.

Motivation & Objective

  • Motivate the importance of assessing omitted variable bias (OVB) under endogenous controls.
  • Develop a design-based framework to compare sensitivity parameters for OVB.
  • Show that residualization-based benchmarks can be biased when controls are endogenous and propose alternatives.
  • Provide identification results and practical metrics (breakdown points) for robustness to unobservables.
  • Illustrate methods with an empirical application and deliver a user-friendly Stata module (regsensitivity).

Proposed method

  • Define the baseline OLS model with observed treatment X, observed controls W1, and unobserved controls W2.
  • Introduce a design distribution framework for covariate observation status and equal selection of covariates.
  • Prove that Oster’s delta parameter can fail to be centered at 1 under endogenous controls (Theorem 3).
  • Introduce two new sensitivity parameters: one comparing selection on observables vs unobservables in the treatment equation, and one in the outcome equation.
  • Derive an identified set for the long regression coefficient (Theorem 4) and provide numerical breakdown points (Theorem 5).
  • Offer exact non-asymptotic distribution results via an empirical DGP and present practical guidance for robustness analysis.
  • Present the companion Stata module regsensitivity for implementation.

Experimental results

Research questions

  • RQ1How does endogenous control affect the benchmarks used in existing sensitivity analyses for omitted variables?
  • RQ2Can we develop sensitivity parameters that compare selection on observables and unobservables without requiring exogenous controls?
  • RQ3What is the appropriate benchmark for equal selection when controls are endogenous, and can we obtain an identifiable robustness metric?
  • RQ4How can researchers compute breakdown points to assess robustness to omitted variables under endogenous controls?
  • RQ5What are the practical implications for empirical studies when applying the new framework to real data?

Key findings

  • Residualization-based sensitivity measures may misrepresent robustness when controls are endogenous (the equal-selection benchmark is not centered at 1 for Oster’s delta).
  • A design-based framework shows that the conventional delta parameter can converge to any real number under equal selection unless controls are exogenous.
  • Two new sensitivity parameters enable comparison of selection on observables and unobservables without requiring exogenous controls and converge to 1 under equal selection (Theorem 2).
  • The identified set for the long regression coefficient is characterized analytically (Theorem 4) and breakdown points can be computed numerically (Theorem 5).
  • Empirical application indicates questionnaire-based outcomes may be sensitive to omitted variables, while election and property tax outcomes remain robust.
  • A Stata module regsensitivity is provided for practical implementation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.