[论文解读] Accuracy of Gaussian approximation in nonparametric Bernstein -- von Mises Theorem
本文在非参数贝叶斯模型中,针对高斯先验,建立了后验集中度与高斯近似精度的非渐近界。结果表明,在温和正则性条件下,后验分布可被以惩罚最大似然估计(pMLE)为中心的高斯分布良好近似,且在中心对称可信集上,误差界为 $ n^{-1} $ 阶,从而通过pMLE的一致性,将贝叶斯可信集与频率学派置信集联系起来。
The prominent Bernstein -- von Mises (BvM) result claims that the posterior distribution after centering by the efficient estimator and standardizing by the square root of the total Fisher information is nearly standard normal. In particular, the prior completely washes out from the asymptotic posterior distribution. This fact is fundamental and justifies the Bayes approach from the frequentist viewpoint. In the nonparametric setup the situation changes dramatically and the impact of prior becomes essential even for the contraction of the posterior; see [vdV2008], [Bo2011], [CaNi2013,CaNi2014] for different models like Gaussian regression or i.i.d. model in different weak topologies. This paper offers another non-asymptotic approach to studying the behavior of the posterior for a special but rather popular and useful class of statistical models and for Gaussian priors. First we derive tight finite sample bounds on posterior contraction in terms of the so called effective dimension of the parameter space. Our main results describe the accuracy of Gaussian approximation of the posterior. In particular, we show that restricting to the class of all centrally symmetric credible sets around pMLE allows to get Gaussian approximation up to order (n^{-1}). We also show that the posterior distribution mimics well the distribution of the penalized maximum likelihood estimator (pMLE) and reduce the question of reliability of credible sets to consistency of the pMLE-based confidence sets. The obtained results are specified for nonparametric log-density estimation and generalized regression.
研究动机与目标
- 为解决在高维或无限维参数空间中经典伯恩斯坦-冯米塞斯(BvM)定理失效的问题,其中先验显著影响后验集中度。
- 以参数空间的有效维数为度量,提供后验集中度与高斯近似误差的非渐近界。
- 建立条件,使贝叶斯可信集可可靠用作频率学派置信集,通过将它们与惩罚最大似然估计(pMLE)的一致性联系起来。
- 分析后验分布的高斯近似精度,特别是当限制在以pMLE为中心的对称可信集时。
提出的方法
- 利用有效维数 $ p_G(\theta) = \mathrm{tr}(H^2 \mathbb{F}_G^{-1}(\theta)) $ 推导后验集中度的有限样本界,其中 $ \mathbb{F}_G $ 为先验下的费雪信息量。
- 引入确保对数似然函数的随机部分为线性且期望对数似然函数为凹的条件,这些条件在高斯回归、对数密度估计和广义线性模型等模型中成立。
- 将惩罚最大似然估计(pMLE)定义为后验众数,其由高斯先验诱导的二次惩罚项决定,并证明后验可模仿pMLE的抽样分布。
- 使用总变差距离量化高斯近似的精度,其界依赖于有效维数及对数似然函数的高阶导数。
- 证明:当可信集被限制在以pMLE为中心的对称区域时,高斯近似误差可提升至 $ O(n^{-1}) $,显著优于一般界。
- 将结果应用于两类关键先验:截断先验(支持在低维子空间上)与平滑先验(先验精度算子特征值衰减),表明对截断先验有 $ p_G \asymp m $,对平滑先验则有效维数随 $ m $ 缓慢增长。
实验结果
研究问题
- RQ1在非参数模型中,当先验为高斯先验时,后验分布在何种条件下可实现 $ O(n^{-1}) $ 阶误差的高斯近似?
- RQ2有效维数 $ p_G(\theta) $ 如何控制后验集中度与高斯近似的精度?
- RQ3贝叶斯可信集在何种条件下可被可靠用作频率学派置信集?
- RQ4先验在高维或非参数模型中对后验的影响程度如何,以及如何量化这种影响?
主要发现
- 当可信集为对称时,后验分布可被以惩罚最大似然估计(pMLE)为中心的高斯分布近似,总变差误差界为 $ O(n^{-1}) $。
- 有效维数 $ p_G(\theta) = \mathrm{tr}(H^2 \mathbb{F}_G^{-1}(\theta)) $ 控制后验集中度速率与高斯近似精度。
- 对于截断先验,当 $ \dim(\mathbb{V}_m) = m $ 时,有 $ p_G(\theta) \leq m $,且当先验方差 $ g_m^2 \ll n $ 时,$ p_G(\theta) \asymp m $。
- 对于平滑先验,有效维数随 $ m $ 缓慢增长,取决于先验精度算子特征值的衰减速率。
- 后验可模仿pMLE的抽样分布,因此可信集的可靠性可归结为pMLE置信集的一致性。
- 在温和正则性条件下结果成立:对数似然函数的随机部分为线性,且期望对数似然函数为凹,这些条件在高斯回归、对数密度估计和广义线性模型等模型中均满足。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。