Skip to main content
QUICK REVIEW

[Paper Review] IV Regressions without Exclusion Restrictions

Wayne Yuan Gao, Rui Wang|arXiv (Cornell University)|Apr 2, 2023
Economic Policies and ImpactsEconomics, Econometrics and Finance3 citations
TL;DR

This paper proposes a novel identification and estimation approach for endogenous linear and nonlinear regression models without excluded instruments, leveraging nonlinear relevance between included exogenous regressors and endogenous variables. By replacing linear first-stage projections with mean projections of the endogenous variable on included exogenous variables, the method achieves point identification under a nonlinearity condition and introduces three estimators—two semiparametric and one discretization-based—showcasing robust finite-sample performance and consistent inference in simulations and empirical applications to returns to education.

ABSTRACT

We study identification and estimation of endogenous linear and nonlinear regression models without excluded instrumental variables, based on the standard mean independence condition and a nonlinear relevance condition. Based on the identification results, we propose two semiparametric estimators as well as a discretization-based estimator that does not require any nonparametric regressions. We establish their asymptotic normality and demonstrate via simulations their robust finite-sample performances with respect to exclusion restrictions violations and endogeneity. Our approach is applied to study the returns to education, and to test the direct effects of college proximity indicators as well as family background variables on the outcome.

Motivation & Objective

  • To address the challenge of identifying endogenous regression models when valid excluded instruments are unavailable.
  • To establish point identification in linear and nonlinear models using only mean independence and a nonlinear relevance condition.
  • To develop feasible, semiparametric and discretization-based estimators that do not rely on nonparametric regressions.
  • To demonstrate robust finite-sample performance under violations of exclusion restrictions and endogeneity.
  • To apply the method to empirical questions on returns to education and direct effects of family background variables.

Proposed method

  • Identification is achieved through the mean projection of the endogenous variable on included exogenous regressors, replacing linear first-stage projections used in standard 2SLS.
  • The key identifying condition is that the conditional expectation function π₀(z) = E[Xᵢ|Zᵢ = z] is nonlinear in z, which is testable using observed data.
  • Two semiparametric estimators are proposed: one based on estimating π₀(z) via SVM or neural networks, and another using a local linear approximation to the structural function.
  • A discretization-based estimator is introduced that avoids nonparametric regressions by creating dummy variables from quantiles of exogenous regressors.
  • Asymptotic normality is established for all three estimators under regularity conditions, enabling valid inference.
  • The approach is extended to nonlinear models with known functional forms, relying on a full-rank condition for local identification.

Experimental results

Research questions

  • RQ1Can endogenous linear regression models be point-identified without excluded instruments?
  • RQ2What conditions on the relationship between included exogenous variables and the endogenous variable enable identification in the absence of exclusion restrictions?
  • RQ3How can robust and efficient estimators be constructed when standard 2SLS fails due to lack of excluded instruments?
  • RQ4How do the proposed estimators perform in finite samples under violations of exclusion restrictions or endogeneity?
  • RQ5What are the empirical implications for estimating returns to education when family background variables may have direct effects?

Key findings

  • The coefficient on education is estimated at 0.109–0.116 (SVM) and 0.109–0.103 (neural network), significantly positive and smaller than 2SLS estimates, suggesting 2SLS overestimates returns to education.
  • Parents’ average education has a significant positive effect on wages (coefficient 0.014–0.029), indicating direct influence beyond education.
  • Number of siblings shows no significant effect on wages across all estimators, with coefficients near zero and p-values above 0.05.
  • The three 2SLS estimators that use parents’ education or siblings as instruments produce higher education returns (0.123–0.149) than the proposed estimators, indicating potential overestimation due to direct effects.
  • The discretization estimator performs comparably to semiparametric estimators and avoids nonparametric regression, offering a computationally simple alternative with good finite-sample properties.
  • Simulations confirm that the proposed estimators maintain good size and power under exclusion restriction violations and endogeneity, outperforming OLS and 2SLS in such settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.