Skip to main content
QUICK REVIEW

[Paper Review] Robust and Efficient Estimation of Potential Outcome Means under Random Assignment

Akanksha Negi, Jeffrey M. Wooldridge|arXiv (Cornell University)|Oct 5, 2020
Economic and Environmental Valuation12 references4 citations
TL;DR

This paper proposes separate linear and nonlinear regression adjustment (RA) methods for estimating potential outcome (PO) means under random assignment with multiple treatment levels. It shows that separate RA is asymptotically at least as efficient as subsample means and generally more efficient than pooled RA unless slopes are homogeneous. Nonlinear RA using quasi-maximum likelihood (e.g., logistic regression) also yields consistent PO mean estimates, with empirical results showing substantial efficiency gains in a contingent valuation study on willingness to pay for oil spill prevention.

ABSTRACT

We study efficiency improvements in randomized experiments for estimating a vector of potential outcome means using regression adjustment (RA) when there are more than two treatment levels. We show that linear RA which estimates separate slopes for each assignment level is never worse, asymptotically, than using the subsample averages. We also show that separate RA improves over pooled RA except in the obvious case where slope parameters in the linear projections are identical across the different assignment levels. We further characterize the class of nonlinear RA methods that preserve consistency of the potential outcome means despite arbitrary misspecification of the conditional mean functions. Finally, we apply these regression adjustment techniques to efficiently estimate the lower bound mean willingness to pay for an oil spill prevention program in California.

Motivation & Objective

  • To improve the efficiency of estimating potential outcome (PO) means when there are more than two treatment levels under random assignment.
  • To extend prior work on average treatment effect estimation to joint estimation of a vector of PO means under multivalued treatments.
  • To characterize nonlinear regression adjustment methods that ensure consistency of PO mean estimators under arbitrary misspecification of conditional mean functions.
  • To evaluate the finite-sample performance of RA methods through simulations and a real-world contingent valuation study.
  • To demonstrate that separate regression adjustment consistently improves precision over subsample means and pooled RA, especially when covariates predict outcomes well.

Proposed method

  • Uses separate linear regression adjustments (SRA) for each treatment level, allowing distinct slope coefficients per group to improve efficiency.
  • Compares SRA to pooled regression adjustment (PRA) and subsample means, showing SRA is never asymptotically worse and is generally more efficient unless slope parameters are identical across groups.
  • Applies quasi-maximum likelihood estimation (QMLE) for nonlinear models (e.g., logistic regression) to ensure consistency of PO mean estimators under model misspecification.
  • Employs simulation studies with three outcome types—fractional, non-negative, and continuous—to assess bias and variance under varying R-squared levels.
  • Applies the methods to a real contingent valuation dataset from Carson et al. (2004) to estimate lower-bound mean willingness to pay (WTP) for oil spill prevention.
  • Uses covariates such as income, tax payment status, proximity to affected area, and environmental attitudes to improve prediction and estimation precision.

Experimental results

Research questions

  • RQ1Is separate linear regression adjustment (SRA) asymptotically more efficient than subsample means for estimating PO means under multivalued random assignment?
  • RQ2Under what conditions does SRA outperform pooled regression adjustment (PRA), and when is the efficiency gain negligible?
  • RQ3Can nonlinear regression adjustment methods (e.g., logistic QMLE) produce consistent estimators of PO means under arbitrary misspecification of the conditional mean function?
  • RQ4How do linear and nonlinear RA methods compare in finite samples, especially when outcomes are fractional or non-negative?
  • RQ5Does regression adjustment improve the precision of lower-bound mean willingness-to-pay (WTP) estimates in a real contingent valuation study?

Key findings

  • Separate linear regression adjustment (SRA) is asymptotically never worse than subsample means and is strictly more efficient unless all slope coefficients are zero.
  • SRA is generally more efficient than pooled RA, except when the true slope parameters are identical across treatment levels, in which case both yield equivalent asymptotic variance.
  • Nonlinear RA using QMLE (e.g., logistic regression) produces consistent estimators of PO means even under model misspecification, provided the mean function is correctly specified within the exponential family.
  • In simulations, both linear and nonlinear RA methods reduce standard errors compared to subsample means, with the magnitude of improvement depending on how well covariates predict potential outcomes.
  • In the empirical application to a California oil spill prevention study, nonlinear pooled logistic regression (pooled RA) produced the smallest standard error for the lower-bound mean WTP estimate, outperforming both linear SRA and nonlinear separate RA.
  • The ABERS estimator (based on subsample means) had the largest standard error, confirming that regression adjustment significantly improves precision in practice.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.