Skip to main content
QUICK REVIEW

[Paper Review] Inferring Treatment Effects After Testing Instrument Strength in Linear Models

Nan Bi, Hyunseung Kang|arXiv (Cornell University)|Mar 14, 2020
Advanced Causal Inference Techniques23 references4 citations
TL;DR

This paper proposes a selective inference method to correct for Type I error inflation in instrumental variables (IV) analysis when treatment effect inference follows a pre-test for instrument strength using an F-statistic. By framing the F-test as a randomized convex optimization problem and applying selective inference techniques, the method provides valid conditional and marginal Type I error control, with reanalysis showing that naive inference can drastically overstate the significance of education's effect on earnings.

ABSTRACT

A common practice in IV studies is to check for instrument strength, i.e. its association to the treatment, with an F-test from regression. If the F-statistic is above some threshold, usually 10, the instrument is deemed to satisfy one of the three core IV assumptions and used to test for the treatment effect. However, in many cases, the inference on the treatment effect does not take into account the strength test done a priori. In this paper, we show that not accounting for this pretest can severely distort the distribution of the test statistic and propose a method to correct this distortion, producing valid inference. A key insight in our method is to frame the F-test as a randomized convex optimization problem and to leverage recent methods in selective inference. We prove that our method provides conditional and marginal Type I error control. We also extend our method to weak instrument settings. We conclude with a reanalysis of studies concerning the effect of education on earning where we show that not accounting for pre-testing can dramatically alter the original conclusion about education's effects.

Motivation & Objective

  • To address the problem of inflated Type I error and invalid p-values that arise when treatment effect inference follows a pre-test for instrument strength in IV models.
  • To develop a method that accounts for the conditioning on the outcome of the pre-test (e.g., F-statistic > 10) to produce valid statistical inference.
  • To extend the method to weak instrument settings using robust test statistics like the CLR test.
  • To demonstrate the practical impact of ignoring pre-testing through re-analysis of landmark studies on education and earnings.
  • To establish a general framework for selective inference in IV settings where diagnostic checks (e.g., instrument strength) are performed before effect estimation.

Proposed method

  • Formulate the pre-test for instrument strength as a randomized convex optimization problem, enabling the use of selective inference tools.
  • Use sampling-based inference to compute the conditional null distribution of the treatment effect test statistic given that the F-test for instrument strength passed (or failed) the threshold.
  • Apply selective inference techniques to derive exact p-values and confidence intervals that are conditionally valid given the pre-test outcome.
  • Extend the method to weak instruments by conditioning on the F-statistic not exceeding the threshold (e.g., C₀ = 10), using weak-instrument-robust test statistics like the CLR test.
  • Implement a simulation-based procedure to generate conditional confidence intervals and p-values that account for the selection event from the pre-test.
  • Ensure both conditional and marginal Type I error control through theoretical proof and empirical validation.

Experimental results

Research questions

  • RQ1How does pre-testing for instrument strength distort the distribution of treatment effect test statistics in IV models?
  • RQ2What is the impact of ignoring the pre-test on Type I error rates and confidence interval coverage in standard IV inference?
  • RQ3Can selective inference techniques be adapted to correct for the conditioning introduced by pre-tests on instrument strength?
  • RQ4How does the proposed method perform in weak instrument settings where traditional IV estimators are biased?
  • RQ5To what extent does pre-testing bias alter conclusions in real-world applications, such as estimating the return to education?

Key findings

  • Naive inference that ignores the pre-test for instrument strength leads to inflated Type I error and confidence intervals with lower-than-nominal coverage, especially when the F-statistic is near the threshold of 10.
  • In the re-analysis of Card's (1995) study on education and earnings, the original p-value of 0.016 using TSLS dropped to 0.602 when pre-testing was properly accounted for, rendering the effect non-significant.
  • For Angrist and Keueger's (1991) study with weak instruments (F-statistic ≈ 0.96), the conditional 95% confidence interval was [-∞, +∞], indicating no precise inference, while the unconditional interval was [0.0238, 0.2671], falsely suggesting precision.
  • The proposed method achieves nominal coverage for confidence intervals and controls both conditional and marginal Type I error rates, even under weak instruments.
  • The method remains effective even when using weak-instrument-robust tests like the CLR test, which are otherwise insensitive to conditioning in standard approaches.
  • The simulation and empirical results demonstrate that failing to account for pre-testing leads to overconfident and misleading conclusions in IV studies, particularly in labor economics applications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.