Skip to main content
QUICK REVIEW

[Paper Review] Statistical Inference with Local Optima

Yen‐Chi Chen|arXiv (Cornell University)|Jul 12, 2018
Statistical Methods and Inference44 references4 citations
TL;DR

This paper studies statistical inference for estimators derived via gradient ascent with multiple random initializations in non-convex, multi-modal likelihood functions. It identifies the population quantity targeted by such estimators, analyzes coverage deficiencies in confidence intervals due to finite initializations, and proposes a two-sample test when the MLE is intractable. The key contribution is a theoretical characterization of the estimator's limiting behavior and confidence interval validity under local optima.

ABSTRACT

We study the statistical properties of an estimator derived by applying a gradient ascent method with multiple initializations to a multi-modal likelihood function. We derive the population quantity that is the target of this estimator and study the properties of confidence intervals (CIs) constructed from asymptotic normality and the bootstrap approach. In particular, we analyze the coverage deficiency due to finite number of random initializations. We also investigate the CIs by inverting the likelihood ratio test, the score test, and the Wald test, and we show that the resulting CIs may be very different. We propose a two-sample test procedure even when the MLE is intractable. In addition, we analyze the performance of the EM algorithm under random initializations and derive the coverage of a CI with a finite number of initializations.

Motivation & Objective

  • To understand the statistical properties of estimators obtained via gradient ascent with multiple random initializations in multi-modal likelihood functions.
  • To identify the population quantity that such estimators target, especially when they converge to local maxima rather than the global MLE.
  • To analyze the coverage properties of confidence intervals constructed via asymptotic normality, bootstrap, and likelihood ratio, score, and Wald tests.
  • To quantify the coverage deficiency in confidence intervals due to a finite number of random initializations.
  • To propose a two-sample test procedure when the MLE is intractable due to non-convexity.

Proposed method

  • Derives the population quantity targeted by gradient ascent estimators with multiple initializations in multi-modal likelihoods using theoretical analysis.
  • Analyzes confidence intervals via asymptotic normality and the bootstrap, showing their coverage may fall short due to finite initialization count.
  • Compares three test-based confidence intervals: likelihood ratio, score, and Wald, demonstrating they can yield significantly different results.
  • Uses concentration inequalities and Hoeffding's inequality to bound the probability of missing the global mode under finite initialization.
  • Introduces event-based analysis (E1, E2, E3, E4) to control error terms in coverage probability, incorporating bias and estimation error in density and gradient estimates.
  • Derives a lower bound on the coverage probability of the confidence interval, showing it is at least $1 - \alpha - (1 - \frac{1}{3}P(\mathcal{A}_{\sf mode}))^n + O(\sqrt{1/(nh^{d+2})}) + O(\sqrt{nh^{d+6}})$.
Figure 1: Log-likelihood function of fitting a 2-Gaussian mixture model to a data that is generated from a 3-Gaussian mixture model. The true distribution has a density function: $p_{0}(x)=0.5\phi(x;0,0.2^{2})+0.45\phi(x;0.75,0.2^{2})+0.05\phi(x;3,0.2^{2})$ , where $\phi(x;\mu,\sigma^{2})$ is the de
Figure 1: Log-likelihood function of fitting a 2-Gaussian mixture model to a data that is generated from a 3-Gaussian mixture model. The true distribution has a density function: $p_{0}(x)=0.5\phi(x;0,0.2^{2})+0.45\phi(x;0.75,0.2^{2})+0.05\phi(x;3,0.2^{2})$ , where $\phi(x;\mu,\sigma^{2})$ is the de

Experimental results

Research questions

  • RQ1What population quantity does a gradient ascent estimator with multiple random initializations target when the likelihood function is multi-modal?
  • RQ2How does a finite number of random initializations affect the coverage probability of confidence intervals constructed via asymptotic normality or bootstrap?
  • RQ3How do confidence intervals based on likelihood ratio, score, and Wald tests compare in terms of coverage when the MLE is not attained?
  • RQ4What is the theoretical coverage guarantee of a confidence interval when the estimator is not the true MLE due to convergence to a local maximum?
  • RQ5Can a two-sample test be constructed when the MLE is intractable due to non-convexity, and what are its theoretical properties?

Key findings

  • The estimator based on multiple random initializations targets a population quantity that is not the global MLE but a local mode that depends on the initialization distribution.
  • Confidence intervals based on asymptotic normality or bootstrap suffer from coverage deficiency due to the finite number of initializations, with coverage probability bounded below by $1 - \alpha - (1 - \frac{1}{3}P(\mathcal{A}_{\sf mode}))^n + O(\sqrt{1/(nh^{d+2})}) + O(\sqrt{nh^{d+6}})$.
  • The likelihood ratio, score, and Wald test-based confidence intervals can yield significantly different coverage, highlighting sensitivity to test choice under local optima.
  • The probability of missing the global mode under finite initialization is bounded by $ (1 - \frac{1}{3}P(\mathcal{A}_{\sf mode}))^n $, which decays exponentially with sample size $n$.
  • The proposed two-sample test is valid even when the MLE is intractable, providing a practical alternative for inference in non-convex models.
  • Theoretical coverage bounds show that coverage deficiency is dominated by the probability of missing the global mode, which is exponentially small under mild regularity conditions.
Figure 3: A modal regression method to the mixture regression problem. Left: we display a data generated by a 4-mixture regression model. Right: the objective function of the linear modal regression as a function of the parameter $\theta$ . The three red vertical lines display the boundary of basins
Figure 3: A modal regression method to the mixture regression problem. Left: we display a data generated by a 4-mixture regression model. Right: the objective function of the linear modal regression as a function of the parameter $\theta$ . The three red vertical lines display the boundary of basins

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.