Skip to main content
QUICK REVIEW

[Paper Review] Statistical models, likelihood, penalized likelihood and hierarchical likelihood

Daniel Commenges|ArXiv.org|Aug 29, 2008
Advanced Statistical Methods and Models26 references3 citations
TL;DR

This paper provides a foundational overview of likelihood inference in statistics, focusing on conventional, penalized, and hierarchical likelihood methods. It establishes their connections to the Kullback-Leibler divergence, demonstrates the equivalence between penalized likelihood and sieve estimators, and explores their relationship to Bayesian MAP estimation, highlighting the challenge of parametrization invariance in such identifications.

ABSTRACT

We give an overview of statistical models and likelihood, together with two of its variants: penalized and hierarchical likelihood. The Kullback-Leibler divergence is referred to repeatedly, for defining the misspecification risk of a model, for grounding the likelihood and the likelihood crossvalidation which can be used for choosing weights in penalized likelihood. Families of penalized likelihood and sieves estimators are shown to be equivalent. The similarity of these likelihood with a posteriori distributions in a Bayesian approach is considered.

Motivation & Objective

  • To clarify the theoretical foundations of statistical models and likelihood, particularly in relation to the Kullback-Leibler divergence.
  • To examine the role of penalized likelihood in introducing smoothness constraints and its equivalence to sieve estimators.
  • To investigate the connections between likelihood-based inference and Bayesian MAP estimation, especially regarding parametrization invariance.
  • To evaluate the use of likelihood cross-validation and AIC-like criteria for model and estimator selection.

Proposed method

  • Uses an extended formulation of the Kullback-Leibler divergence that explicitly accounts for sigma-fields, linking it to model misspecification risk.
  • Defines likelihood and penalized likelihood through the lens of Radon-Nikodym derivatives and conditional expectations.
  • Demonstrates equivalence between families of penalized likelihood estimators and sieve estimators via functional approximation.
  • Applies likelihood cross-validation and AIC-type criteria to estimate Kullback-Leibler risk, replacing expectation under true distribution with empirical expectation.
  • Explores the identification of maximum likelihood and penalized likelihood estimators with MAP estimators under specific priors, particularly flat and Jeffreys' priors.
  • Highlights the non-invariance of MAP estimators under reparametrization, challenging the direct equivalence between penalized likelihood and Bayesian MAP.

Experimental results

Research questions

  • RQ1How does the Kullback-Leibler divergence underlie the risk assessment of statistical models and likelihood-based estimators?
  • RQ2What is the mathematical and conceptual equivalence between penalized likelihood and sieve estimators?
  • RQ3Can maximum penalized likelihood estimators be interpreted as MAP estimators, and what are the limitations of this interpretation?
  • RQ4How can likelihood cross-validation be used to select optimal estimators or models based on Kullback-Leibler risk?
  • RQ5What are the implications of parametrization dependence for identifying likelihood-based estimators with Bayesian MAP estimators?

Key findings

  • Penalized likelihood estimators and sieve estimators form equivalent families, establishing a theoretical bridge between nonparametric smoothing and regularization.
  • The Kullback-Leibler divergence serves as the fundamental risk measure underlying maximum likelihood estimation, AIC, and likelihood cross-validation.
  • Likelihood cross-validation can estimate the Kullback-Leibler risk by replacing the true expectation with the empirical distribution, enabling practical model selection.
  • Maximum penalized likelihood estimators can be interpreted as MAP estimators under specific priors, but this equivalence breaks under reparametrization due to non-invariant priors.
  • Jeffreys' prior, while invariant under reparametrization, shifts the MAP estimate away from the MLE, showing that flat priors and MLE are not equivalent under change of parameterization.
  • The paper concludes that while likelihood and Bayesian MAP estimators are closely related, the lack of invariance under reparametrization limits the strength of their identification.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.