[论文解读] High confidence estimates of the mean of heavy-tailed real random variables
本文提出了一种用于重尾分布均值的PAC-Bayesian迭代截断估计器,实现了接近极小极大最优水平的非渐近置信区间,其区间宽度与高斯经验均值的偏差界相当。该方法在方差或峰度有界的情况下仍能实现高置信度估计,在最坏情况的重尾场景下优于经验均值。
We present new estimators of the mean of a real valued random variable, based on PAC-Bayesian iterative truncation. We analyze the non-asymptotic minimax properties of the deviations of estimators for distributions having either a bounded variance or a bounded kurtosis. It turns out that these minimax deviations are of the same order as the deviations of the empirical mean estimator of a Gaussian distribution. Nevertheless, the empirical mean itself performs poorly at high confidence levels for the worst distribution with a given variance or kurtosis (which turns out to be heavy tailed). To obtain (nearly) minimax deviations in these broad class of distributions, it is necessary to use some more robust estimator, and we describe an iterated truncation scheme whose deviations are close to minimax. In order to calibrate the truncation and obtain explicit confidence intervals, it is necessary to dispose of a prior bound either on the variance or the kurtosis. When a prior bound on the kurtosis is available, we obtain as a by-product a new variance estimator with good large deviation properties. When no prior bound is available, it is still possible to use Lepski's approach to adapt to the unknown variance, although it is no more possible to obtain observable confidence intervals.
研究动机与目标
- 为具有重尾分布的实值随机变量的均值构建非渐近置信区间。
- 在估计中实现高置信水平(例如,1−ε,其中ε极小),特别是在模型选择等多重比较场景下。
- 构造偏差界接近极小极大的估计器,适用于方差或峰度有界的分布类。
- 通过利用方差或峰度的先验界,实现可观测的置信区间,并作为副产品推导出一种新的鲁棒方差估计器。
- 证明在最坏情况的重尾分布下,经验均值在高置信水平下表现不佳,从而凸显对鲁棒替代方法的需求。
提出的方法
- 提出一种基于分段线性有界损失函数 $ L(x) $ 的截断均值估计器,其中 $ L_{-}(x) riangleq -L_{+}(-x) $,旨在限制极端值的影响。
- 采用PAC-Bayesian框架,推导估计器偏离真实均值的矩生成函数界,确保精确的尾部控制。
- 引入一种迭代截断方案,通过校准的截断参数 $ heta_0 $ 递归更新估计器,提升鲁棒性。
- 使用校准参数 $ heta_0 $ 和调参参数 $ eta $,其中 $ eta = rac{1}{ heta_0} $,以控制截断水平并推导置信区间。
- 推导出形式为 $ igl| heta_{ ext{est}} - migr| riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{a heta_0^2}{2}(v + (m - heta_0)^2)) + ext{log}(rac{1}{ ho}) igr] $ 的置信区间,其中 $ a riangleq rac{2[ ext{exp}( heta_0) - 1 - heta_0]}{ heta_0^2} riangleq 1.2 $。
- 通过 $ heta_0 = rac{1}{ heta_0} = rac{1}{ heta_0} $ 校准截断参数 $ heta_0 $,最终得到偏差界 $ igl| heta_{ ext{est}} - migr| riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{a heta_0^2}{2}(v + (m - heta_0)^2)) + ext{log}(rac{1}{ ho}) igr] $,其中 $ heta_0 $ 选择为使界最小化。
实验结果
研究问题
- RQ1我们能否在不依赖高斯假设的前提下,为重尾分布的均值构造一个保持高置信水平(例如,1−ε,其中ε极小)的非渐近置信区间?
- RQ2在方差或峰度有界的分布类中,经验均值的偏差界与极小极大最优界相比如何?
- RQ3能否设计一种鲁棒估计器,使其偏差界在真实分布为重尾时仍接近极小极大?
- RQ4与最优截断函数相比,使用更简单的截断函数(例如,分段线性)在偏差界膨胀方面的代价是什么?
- RQ5在何种条件下,我们可以在不知道方差或峰度的情况下构造可观测的置信区间,以及何时需要使用Lepski方法?
主要发现
- 即使在方差有界的情况下,经验均值估计器在最坏情况的重尾分布下于高置信水平下表现极差。
- 所提出的PAC-Bayesian迭代截断估计器的偏差界仅比极小极大最优界低约10%——具体为 $ igl| heta_{ ext{est}} - migr| riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{a heta_0^2}{2}(v + (m - heta_0)^2)) + ext{log}(rac{1}{ ho}) igr] $,其中 $ a riangleq 1.2 $,最终界为 $ riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{1.2 heta_0^2}{2}(v + (m - heta_0)^2)) + ext{log}(rac{1}{ ho}) igr] $。
- 当存在先验界 $ v_0 $ 和 $ heta_0 $ 时,估计器的置信区间宽度为 $ riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{1.2 heta_0^2}{2}(v_0 + heta_0^2)) + ext{log}(rac{1}{ ho}) igr] $,其中 $ heta_0 $ 选择为 $ heta_0 = rac{1}{ heta_0} = rac{1}{ heta_0} $,最终偏差界为 $ riangleq rac{1}{ heta_0} igl[ ext{log}(1 + heta_0(m - heta_0) + rac{1.2 heta_0^2}{2}(v_0 + heta_0^2)) + ext{log}(rac{1}{ ho}) igr] $。
- 该估计器的偏差在极小极大最优界的一个因子 $ 1.1 $ 之内,对应于与高斯经验均值相比约10%的精度损失,即使在重尾分布下亦成立。
- 当存在峰度的先验界时,该方法可导出一种具有强大偏差性质的新方差估计器。
- 在缺乏对方差或峰度的先验知识时,可使用Lepski方法自适应未知尺度,但若无额外假设,则无法构造可观测的置信区间。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。