Skip to main content
QUICK REVIEW

[Paper Review] Estimate-Then-Optimize versus Integrated-Estimation-Optimization versus Sample Average Approximation: A Stochastic Dominance Perspective

Adam N. Elmachtoub, Henry Lam|arXiv (Cornell University)|Apr 13, 2023
Advanced Bandit Algorithms Research4 citations
TL;DR

This paper compares estimate-then-optimize (ETO), integrated-estimation-optimization (IEO), and sample average approximation (SAA) in data-driven stochastic optimization. Using first-order stochastic dominance of regret, it shows that when the model is well-specified and data is sufficient, ETO asymptotically dominates IEO, which in turn dominates SAA—reversing the typical performance order under model misspecification.

ABSTRACT

In data-driven stochastic optimization, model parameters of the underlying distribution need to be estimated from data in addition to the optimization task. Recent literature considers integrating the estimation and optimization processes by selecting model parameters that lead to the best empirical objective performance. This integrated approach, which we call integrated-estimation-optimization (IEO), can be readily shown to outperform simple estimate-then-optimize (ETO) when the model is misspecified. In this paper, we show that a reverse behavior appears when the model class is well-specified and there is sufficient data. Specifically, for a general class of nonlinear stochastic optimization problems, we show that simple ETO outperforms IEO asymptotically when the model class covers the ground truth, in the strong sense of stochastic dominance of the regret. Namely, the entire distribution of the regret, not only its mean or other moments, is always better for ETO compared to IEO. Our results also apply to constrained, contextual optimization problems where the decision depends on observed features. Whenever applicable, we also demonstrate how standard sample average approximation (SAA) performs the worst when the model class is well-specified in terms of regret, and best when it is misspecified. Finally, we provide experimental results to support our theoretical comparisons and illustrate when our insights hold in finite-sample regimes and under various degrees of misspecification.

Motivation & Objective

  • To theoretically compare ETO, IEO, and SAA in data-driven stochastic optimization under well-specified models.
  • To investigate whether IEO's computational cost is justified by performance gains when the model class contains the true distribution.
  • To analyze the asymptotic regret distribution of each method using stochastic dominance, focusing on solution quality beyond moments.
  • To clarify the conditions under which simpler ETO can outperform more complex IEO, especially in contextual and constrained optimization.
  • To provide a principled statistical framework for comparing data-driven optimization methods at the distribution level of regret.

Proposed method

  • Uses first-order stochastic dominance to compare the full distribution of regret across ETO, IEO, and SAA, rather than just mean or variance.
  • Analyzes regret as the optimality gap between the data-driven solution and the true optimal solution under the true distribution.
  • Applies asymptotic analysis under the assumption that the model class contains the true distribution and sufficient data is available.
  • Derives theoretical conditions under which ETO stochastically dominates IEO in regret distribution, even though IEO is designed to optimize empirical performance.
  • Extends results to constrained and contextual optimization problems, requiring additional assumptions on constraint qualifications and asymptotic normality.
  • Compares SAA to ETO and IEO, showing SAA performs worst under well-specified models but best under misspecification.

Experimental results

Research questions

  • RQ1Under what conditions does ETO stochastically dominate IEO in terms of regret distribution when the model is well-specified?
  • RQ2How does the performance ranking of ETO, IEO, and SAA change between well-specified and misspecified models?
  • RQ3Does the integrated-estimation-optimization (IEO) approach provide a meaningful advantage over ETO in large-sample, well-specified settings?
  • RQ4How does sample average approximation (SAA) compare to ETO and IEO in terms of regret distribution under model well-specification?
  • RQ5What are the implications of these regret-based comparisons for choosing between computationally simpler ETO and more complex IEO in practice?

Key findings

  • When the model class covers the true distribution and data is sufficient, ETO asymptotically stochastically dominates IEO in terms of regret distribution.
  • IEO stochastically dominates SAA under well-specified models, reversing the typical performance ranking where SAA is often preferred.
  • The stochastic dominance is strong—meaning the entire cumulative distribution of regret is better for ETO than IEO—implying better generalization across all quantiles.
  • This dominance holds for both unconstrained and constrained, contextual optimization problems under standard regularity conditions.
  • In finite samples, the theoretical dominance of ETO over IEO holds under moderate misspecification, but breaks down as misspecification increases.
  • SAA performs worst in terms of regret under well-specified models but becomes optimal under model misspecification, highlighting a fundamental trade-off.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.