[论文解读] Sharp regret bounds for empirical Bayes and compound decision problems
本文为正态和泊松位置模型中的经验贝叶斯与复合决策问题建立了精确的渐近后悔界限。证明了在紧支集先验下,最优后悔的量级为 $Θ((\log n / \log\log n)^2)$,在子指数先验下为 $Θ(\log^3 n)$,并给出了在正态模型中相应的下界 $Ω((\log n / \log\log n)^2)$ 和 $Ω(\log^2 n)$,从而解决了长期存在的猜想,并确立了数据驱动收缩估计器的紧致性能极限。
We consider the classical problems of estimating the mean of an $n$-dimensional normally (with identity covariance matrix) or Poisson distributed vector under the squared loss. In a Bayesian setting the optimal estimator is given by the prior-dependent conditional mean. In a frequentist setting various shrinkage methods were developed over the last century. The framework of empirical Bayes, put forth by Robbins (1956), combines Bayesian and frequentist mindsets by postulating that the parameters are independent but with an unknown prior and aims to use a fully data-driven estimator to compete with the Bayesian oracle that knows the true prior. The central figure of merit is the regret, namely, the total excess risk over the Bayes risk in the worst case (over the priors). Although this paradigm was introduced more than 60 years ago, little is known about the asymptotic scaling of the optimal regret in the nonparametric setting. We show that for the Poisson model with compactly supported and subexponential priors, the optimal regret scales as $Θ((\frac{\log n}{\log\log n})^2)$ and $Θ(\log^3 n)$, respectively, both attained by the original estimator of Robbins. For the normal mean model, the regret is shown to be at least $Ω((\frac{\log n}{\log\log n})^2)$ and $Ω(\log^2 n)$ for compactly supported and subgaussian priors, respectively, the former of which resolves the conjecture of Singh (1979) on the impossibility of achieving bounded regret; before this work, the best regret lower bound was $Ω(1)$. In addition to the empirical Bayes setting, these results are shown to hold in the compound setting where the parameters are deterministic. As a side application, the construction in this paper also leads to improved or new lower bounds for density estimation of Gaussian and Poisson mixtures.
研究动机与目标
- 为非参数设定下的经验贝叶斯与复合决策问题建立精确的渐近后悔界限。
- 解决Singh(1979)关于在紧支集先验下正态均值模型中无法实现有界后悔的猜想。
- 在各种先验假设(紧支集、次高斯、次指数)下,推导正态和泊松模型中后悔的紧致上下界。
- 将结果扩展到参数为确定性而非随机的复合决策设定,表明后悔量级具有等价性。
- 将构造方法应用于推导高斯和泊松位置尺度混合密度估计的新或改进的极小极大下界。
提出的方法
- 将后悔分析简化为通过一般化后悔下界程序估计回归函数。
- 使用阿苏勒引理和广义费诺方法的变体,通过精心构造的参数配置先验,推导出后悔的下界。
- 采用截断技术控制重尾或无界先验的影响,表明后悔对截断具有鲁棒性,误差在可控范围内。
- 应用罗宾斯估计器(一种经典的经脸贝叶斯方法),证明其在正态和泊松模型中均能达到最优的后悔量级。
- 利用积分算子的谱性质及正交基函数(如埃尔米特和拉盖尔函数)构造具有可控矩和范数的测试先验。
- 利用数据处理不等式和赫林格距离界,将混合密度估计误差与决策问题中的后悔联系起来。
实验结果
研究问题
- RQ1在紧支集先验下,经验贝叶斯框架中正态均值模型的最优渐近后悔量级为何?
- RQ2Singh(1979)关于在正态均值模型中无法实现有界后悔的猜想是否可被解决?其在正态均值模型中的真实后悔下界为何?
- RQ3在紧支集和次指数先验下,泊松模型中的精确后悔量级为何?
- RQ4经验贝叶斯设定下的后悔界限与参数为确定性的复合决策设定下的后悔界限有何比较?
- RQ5用于后悔下界构造的方法是否可被重新用于推导高斯和泊松混合密度估计的新极小极大下界?
主要发现
- 在泊松模型中,对于紧支集先验,最优后悔量级为 $Θ((\log n / \log\log n)^2)$;对于次指数先验,量级为 $Θ(\log^3 n)$,均由罗宾斯估计器实现。
- 在正态均值模型中,对于紧支集先验,后悔至少为 $Ω((\log n / \log\log n)^2)$,从而解决了Singh(1979)关于有界后悔不可能实现的猜想。
- 对于正态模型中的次高斯先验,后悔至少为 $Ω(\log^2 n)$,优于此前最佳下界 $Ω(1)$。
- 在复合决策设定中,相同的后悔量级成立,其中参数为确定性而非随机,表明根本性能极限具有等价性。
- 该构造导出了高斯混合密度估计的新极小极大下界:紧支集均值下界为 $Ω(\frac{\log n}{n \log\log n})$,次高斯混合下界为 $Ω(\frac{\log n}{n})$。
- 结果是紧致的,因为下界与罗宾斯估计器在各自模型中实现的上界完全匹配。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。