[论文解读] Minimax Estimation of KL Divergence between Discrete Distributions.
本文提出了一种在大字母表设置下,针对两个离散分布之间KL散度的极小化最大率最优估计器,利用先进的近似方法,实现了与MLE使用n ln n个样本时相当的性能。该估计器具有自适应性,无需事先知道支撑集大小或似然比边界,且表明有效样本量扩大效应的适用范围比以往所知更广。
We consider the problem of estimating the KL divergence between two discrete probability measures $P$ and $Q$ from empirical data in a non-asymptotic and possibly large alphabet setting. We construct minimax rate-optimal estimators for $D(P\|Q)$ when the likelihood ratio is upper bounded by a constant which may depend on the support size, and show that the performance of the optimal estimator with $n$ samples is essentially that of the Maximum Likelihood Estimator (MLE) with $n\ln n$ samples. Our estimator is adaptive in the sense that it does not require the knowledge of the support size or the upper bound on the likelihood ratio. Our approach refines the \emph{Approximation} methodology recently developed for the construction of near minimax estimators of functionals of high-dimensional parameters, such as entropy, R\'enyi entropy, mutual information and $\ell_1$ distance in large alphabet settings, and shows that the \emph{effective sample size enlargement} phenomenon holds significantly more widely than previously established.
研究动机与目标
- 开发一种在大字母表设置下,针对离散分布之间KL散度的非渐近、极小化最大率最优估计器。
- 确定在何种条件下KL估计的有效样本量可扩大ln n倍,从而实现与使用n ln n个样本的MLE性能相当。
- 构建一种可自适应未知支撑集大小和似然比边界、且无需先验知识的估计器。
- 将近似方法的适用范围扩展至KL散度,确认其在熵和互信息之外的更广泛应用性。
提出的方法
- 该方法采用一种为高维泛函量专门设计的精细化近似框架,特别适配KL散度估计。
- 引入一种偏差校正估计器,利用在有界性约束下的经验似然比。
- 通过非渐近分析推导极小化最大风险界,确保在该类分布中所有分布下均实现最优性能。
- 通过避免对支撑集大小或似然比上界显式依赖,使估计器具备自适应性。
- 利用大数定律和经验过程技术,推导出理论保证,以控制估计误差。
- 该框架推广了先前基于近似的估计方法,将其扩展至KL散度,并证明了其最优性。
实验结果
研究问题
- RQ1能否在非渐近、大字母表设置下,构建KL散度的极小化最大率最优估计器?
- RQ2此前在熵和互信息中观察到的有效样本量扩大现象——即样本量扩大ln n倍——在KL散度中在多大程度上仍然成立?
- RQ3此类估计器能否对未知的支撑集大小和似然比边界实现自适应?
- RQ4此前用于其他泛函量的近似方法,是否也能在KL散度中产生最优结果?
主要发现
- 所提出的估计器在有界似然比约束下,实现了KL散度的极小化最大风险最优性。
- 即使仅使用n个样本,该估计器的性能在渐近意义上等价于使用n ln n个样本的MLE。
- 该方法具有自适应性,无需知晓支撑集大小或似然比上界。
- 有效样本量扩大现象在KL散度中依然成立,证实其适用范围不仅限于熵和互信息。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。