Skip to main content
QUICK REVIEW

[Paper Review] Bayesian Posteriors Without Bayes' Theorem

Theodore P. Hill, Marco Dall’Aglio|arXiv (Cornell University)|Mar 1, 2012
Bayesian Modeling and Causal Inference5 references3 citations
TL;DR

This paper demonstrates that the Bayesian posterior arises naturally as the unique solution to multiple optimization problems—minimizing loss of Shannon information, minimizing the maximin likelihood ratio, and ensuring proportional consolidation—without requiring Bayes' Theorem or conditional probability interpretations. The key contribution is a foundational justification of Bayesian posteriors through information-theoretic and optimization principles, offering a reconciliation between classical and Bayesian statistics.

ABSTRACT

The classical Bayesian posterior arises naturally as the unique solution of several different optimization problems, without the necessity of interpreting data as conditional probabilities and then using Bayes' Theorem. For example, the classical Bayesian posterior is the unique posterior that minimizes the loss of Shannon information in combining the prior and the likelihood distributions. These results, direct corollaries of recent results about conflations of probability distributions, reinforce the use of Bayesian posteriors, and may help partially reconcile some of the differences between classical and Bayesian statistics.

Motivation & Objective

  • To establish Bayesian posteriors as the unique solution to optimization problems involving information loss, without relying on conditional probability or Bayes' Theorem.
  • To reconcile classical and Bayesian statistical frameworks by showing that Bayesian posteriors emerge naturally from information-theoretic optimization.
  • To generalize the posterior distribution to weighted prior and likelihood distributions, where prior and likelihood are assigned different importance.
  • To identify the optimal posterior as the one minimizing maximum loss of weighted Shannon information, extending the classical Bayesian framework.

Proposed method

  • Formalizes the combined Shannon information of prior and likelihood as the sum of their individual surprisal values: $ S_{\{P_0,L\}}(A) = -\log_2 P_0(A)L(A) $.
  • Defines the loss of information as the maximum difference between combined and posterior surprisal: $ M(P_1; P_0, L) = \max_A \left\{ \log_2 \frac{P_1(A)}{P_0(A)L(A)} \right\} $.
  • Identifies the Bayesian posterior as the unique distribution minimizing this maximum loss of Shannon information.
  • Introduces a weighted version of Shannon information using relative weights $ w_0 $ and $ w_L $, normalizing by $ \max(w_0, w_L) $.
  • Derives the optimal weighted posterior as $ p^w_1(\theta) \propto (p_0(\theta))^{w_0 / \max(w_0,w_L)} (p_L(\theta))^{w_L / \max(w_0,w_L)} $, which reduces to the standard Bayesian posterior when $ w_0 = w_L $.
  • Establishes equivalence of the Bayesian posterior to solutions of minimax likelihood ratio and proportional consolidation criteria.

Experimental results

Research questions

  • RQ1Can the Bayesian posterior be derived without invoking Bayes' Theorem or conditional probability?
  • RQ2Is the Bayesian posterior the unique solution to an optimization problem minimizing loss of Shannon information in combining prior and likelihood?
  • RQ3How can the Bayesian posterior be generalized when prior and likelihood are assigned different weights?
  • RQ4Does the Bayesian posterior also minimize the maximin likelihood ratio between prior and likelihood distributions?
  • RQ5Is the Bayesian posterior the unique proportional consolidation of prior and likelihood distributions?

Key findings

  • The Bayesian posterior is the unique distribution that minimizes the maximum loss of Shannon information when combining prior and likelihood distributions.
  • The Bayesian posterior also minimizes the maximin likelihood ratio between the prior and likelihood, confirming its optimality under a different criterion.
  • The Bayesian posterior is the unique proportional consolidation of prior and likelihood, meaning it preserves relative odds between outcomes.
  • When prior and likelihood are assigned different weights $ w_0 $ and $ w_L $, the optimal posterior is given by a power-weighted product of the prior and likelihood densities, normalized.
  • The weighted posterior reduces to the standard Bayesian posterior when $ w_0 = w_L $, confirming consistency with classical Bayesian updating.
  • The results provide a non-Bayesian foundation for Bayesian posteriors using information-theoretic and optimization principles, supporting their use across statistical paradigms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.