Skip to main content
QUICK REVIEW

[论文解读] Bayesian Entropy Estimation for Countable Discrete Distributions

Evan Archer, Il Memming Park|arXiv (Cornell University)|Feb 2, 2013
Bayesian Methods and Mixture Models参考文献 30被引用 6
一句话总结

本文提出了一种基于Pitman-Yor混合(PYM)先验的贝叶斯熵估计器,用于可数无限的离散分布,该先验在熵上诱导出近乎非信息性的平坦先验。通过在连续混合测度下混合Pitman-Yor过程,该方法在样本不足的区域实现了稳定的熵估计,其性能在模拟数据和真实神经数据中均优于固定先验方法,且熵估计的后验矩具有解析可计算性。

ABSTRACT

We consider the problem of estimating Shannon's entropy $H$ from discrete data, in cases where the number of possible symbols is unknown or even countably infinite. The Pitman-Yor process, a generalization of Dirichlet process, provides a tractable prior distribution over the space of countably infinite discrete distributions, and has found major applications in Bayesian non-parametric statistics and machine learning. Here we show that it also provides a natural family of priors for Bayesian entropy estimation, due to the fact that moments of the induced posterior distribution over $H$ can be computed analytically. We derive formulas for the posterior mean (Bayes' least squares estimate) and variance under Dirichlet and Pitman-Yor process priors. Moreover, we show that a fixed Dirichlet or Pitman-Yor process prior implies a narrow prior distribution over $H$, meaning the prior strongly determines the entropy estimate in the under-sampled regime. We derive a family of continuous mixing measures such that the resulting mixture of Pitman-Yor processes produces an approximately flat prior over $H$. We show that the resulting Pitman-Yor Mixture (PYM) entropy estimator is consistent for a large class of distributions. We explore the theoretical properties of the resulting estimator, and show that it performs well both in simulation and in application to real data.

研究动机与目标

  • 解决在可能符号数量未知或可数无限时,从稀疏、样本不足的数据中估计香农熵的挑战。
  • 克服在低样本区域中,固定狄利克雷或Pitman-Yor过程先验导致的偏差和可信区间过窄的问题。
  • 构建一种针对离散分布的先验,使其在熵上诱导出近乎平坦的、非信息性的先验,从而实现稳健的贝叶斯推断。
  • 推导在所提先验下熵估计后验均值和方差的解析表达式。
  • 在模拟数据和真实神经数据上,展示所得到的Pitman-Yor混合(PYM)估计器的一致性和优异性能。

提出的方法

  • 将Pitman-Yor过程(PYP)用作无限维离散分布的先验,其推广了狄利克雷过程,可允许幂律尾部行为。
  • 推导了在狄利克雷和PYP先验下熵的后验均值和方差的解析表达式,从而实现高效的贝叶斯估计。
  • 引入PYP超参数的连续混合测度族,以使熵上的边际先验分布趋于平坦,从而形成Pitman-Yor混合(PYM)先验。
  • 利用Beta函数和Gamma函数的渐近展开,分析在样本量增加时后验矩的极限行为。
  • 应用关于边缘似然(证据)在折扣参数$d$和聚集参数$α$上的单峰性定理,以确保推断的稳定性。
  • 通过模拟和视网膜记录的真实神经数据验证该估计器,与标准估计器进行性能比较。

实验结果

研究问题

  • RQ1能否为支持未知或无限的离散分布构建一种在熵上具有非信息性先验的贝叶斯方法?
  • RQ2在样本不足的区域中,固定PYP或狄利克雷过程先验如何影响熵估计?它们会引入何种偏差?
  • RQ3能否构造一个PYP的混合,使得其在熵上的边际先验近似平坦?
  • RQ4在所提PYM先验下,熵的后验均值和方差具有何种解析性质?
  • RQ5所得到的PYM熵估计器是否对一大类基础分布具有的一致性?

主要发现

  • 在狄利克雷和Pitman-Yor过程先验下,熵的后验均值和方差可进行解析计算,从而实现高效的贝叶斯估计。
  • 固定PYP或狄利克雷先验在熵上诱导出狭窄的先验分布,导致在低样本区域出现强烈偏差和过度自信的可信区间。
  • 通过连续混合测度构建的所提Pitman-Yor混合(PYM)先验,使熵上的先验分布近似平坦,从而减少了先验引入的偏差。
  • 即使真实支持是有限但未知的,PYM估计器对一大类离散分布仍具有一致性。
  • 在模拟和真实神经数据中,PYM估计器在偏差和均方误差方面均优于标准的频率学派和贝叶斯估计器。
  • 边缘似然(证据)在折扣参数$d$和聚集参数$α$上均为单峰,确保了优化和推断的稳定性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。