Skip to main content
QUICK REVIEW

[論文レビュー] High confidence estimates of the mean of heavy-tailed real random variables

Olivier Catoni|ArXiv.org|Sep 29, 2009
Statistical Methods and Inference参考文献 5被引用数 5
ひとこと要約

この論文は、重たい尾を持つ分布の平均に対するPAC-Bayesian反復型截断推定量を導入し、漸近的でない信頼区間を実現した。その幅は最小最大最適水準に近く、ガウス分布の経験的平均が達成する乖離境界と一致する。この手法により、分散が有界または尖度が有界な状況下でも高い信頼性の推定が可能となり、最悪の重たい尾を持つ状況下で経験的平均を上回る性能を発揮する。

ABSTRACT

We present new estimators of the mean of a real valued random variable, based on PAC-Bayesian iterative truncation. We analyze the non-asymptotic minimax properties of the deviations of estimators for distributions having either a bounded variance or a bounded kurtosis. It turns out that these minimax deviations are of the same order as the deviations of the empirical mean estimator of a Gaussian distribution. Nevertheless, the empirical mean itself performs poorly at high confidence levels for the worst distribution with a given variance or kurtosis (which turns out to be heavy tailed). To obtain (nearly) minimax deviations in these broad class of distributions, it is necessary to use some more robust estimator, and we describe an iterated truncation scheme whose deviations are close to minimax. In order to calibrate the truncation and obtain explicit confidence intervals, it is necessary to dispose of a prior bound either on the variance or the kurtosis. When a prior bound on the kurtosis is available, we obtain as a by-product a new variance estimator with good large deviation properties. When no prior bound is available, it is still possible to use Lepski's approach to adapt to the unknown variance, although it is no more possible to obtain observable confidence intervals.

研究の動機と目的

  • 実数値の確率変数が重たい尾を持つ分布に従う場合の平均に対する漸近的でない信頼区間の構築を目的とする。
  • 特に、多数の比較を伴うモデル選択のような状況において、非常に小さなεに対する1−εのような高い信頼水準での推定を達成することを目的とする。
  • 分散が有界または尖度が有界な分布のクラスにおいて、ほぼ最小最大最適な乖離境界を持つ推定量の構築を目的とする。
  • 分散または尖度の事前境界を活用することで観測可能な信頼区間を構築し、同時に新たなロバストな分散推定量を副産物として得ることを目的とする。
  • 経験的平均が最悪の重たい尾を持つ分布下で高い信頼水準において失敗することを示し、これによりロバストな代替手法の必要性を示すこと。

提案手法

  • 極端な値の影響を制限するために、区分的線形かつ有界な損失関数$L(x)$に基づく截断平均推定量を提案。ここで$L_{-}(x) \triangleq -L_{+}(-x)$である。
  • PAC-Bayesian枠組みを用いて、推定量の真の平均からの乖離に関する指数モーメントの境界を導出し、鋭い尾部制御を実現する。
  • 適切にキャリブレーションされた截断パラメータ$\theta_0$を用いて、再帰的に更新される反復的截断スキームを導入し、より高いロバスト性を実現する。
  • 截断レベルと信頼区間の導出を制御するため、$\theta_0$と調整パラメータ$\beta$を$\beta = \frac{1}{\theta_0}$と定義する。
  • 信頼区間の形式$\bigl|\theta_{\text{est}} - m\bigr| \triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{a\theta_0^2}{2}(v + (m - \theta_0)^2)) + \text{log}(\frac{1}{\rho}) \bigr]$を導出。ここで$a \triangleq \frac{2[\text{exp}(\theta_0) - 1 - \theta_0]}{\theta_0^2} \triangleq 1.2$である。
  • 截断パラメータ$\theta_0$を$\theta_0 = \frac{1}{\theta_0} = \frac{1}{\theta_0}$としてキャリブレーションし、最終的な乖離境界を$\bigl|\theta_{\text{est}} - m\bigr| \triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{a\theta_0^2}{2}(v + (m - \theta_0)^2)) + \text{log}(\frac{1}{\rho}) \bigr]$とする。$\theta_0$は境界を最小化するように選ぶ。

実験結果

リサーチクエスチョン

  • RQ1重たい尾を持つ分布の平均に対して、ガウス分布の仮定に依存せずに、非常に高い信頼水準(例:εが非常に小さい場合の1−ε)を維持する漸近的でない信頼区間を構築できるか?
  • RQ2分散が有界または尖度が有界な分布のクラスにおいて、経験的平均の乖離境界は最小最大最適境界と比べてどの程度異なるか?
  • RQ3真の分布が重たい尾を持つ場合でも、その乖離境界がほぼ最小最大最適となるようなロバストな推定量を設計できるか?
  • RQ4最適な損失関数と比較して、より単純な截断関数(例:区分的線形)を用いる場合、乖離境界の増幅はどの程度になるか?
  • RQ5分散または尖度の事前知識がない状況で、観測可能な信頼区間を構築できる条件は何か?また、Lepskiの手法が必要となる状況は何か?

主な発見

  • 分散が有界であっても、最悪の重たい尾を持つ分布下では経験的平均推定量は高い信頼水準において著しく性能を発揮しない。
  • 提案されたPAC-Bayesian反復型截断推定量は、最小最大最適境界の10%以内の乖離境界を達成する。具体的には、$\bigl|\theta_{\text{est}} - m\bigr| \triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{a\theta_0^2}{2}(v + (m - \theta_0)^2)) + \text{log}(\frac{1}{\rho}) \bigr]$、$a \triangleq 1.2$であり、境界は$\triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{1.2\theta_0^2}{2}(v + (m - \theta_0)^2)) + \text{log}(\frac{1}{\rho}) \bigr]$となる。
  • 事前境界$v_0$と$\theta_0$が利用可能な場合、推定量は信頼区間幅$\triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{1.2\theta_0^2}{2}(v_0 + \theta_0^2)) + \text{log}(\frac{1}{\rho}) \bigr]$を達成する。ここで$\theta_0$は$\theta_0 = \frac{1}{\theta_0} = \frac{1}{\theta_0}$として選ばれる。最終的な乖離境界は$\triangleq \frac{1}{\theta_0} \bigl[ \text{log}(1 + \theta_0(m - \theta_0) + \frac{1.2\theta_0^2}{2}(v_0 + \theta_0^2)) + \text{log}(\frac{1}{\rho}) \bigr]$となる。
  • 推定量の乖離は最小最大最適境界の1.1倍以内に収まり、ガウス分布の経験的平均と比較して10%の精度損失にとどまる。これは重たい尾を持つ分布下でも同様に成立する。
  • 尖度の事前境界が利用可能な場合、本手法は強力な大偏差性質を有する新たな分散推定量を導出する。
  • 分散または尖度の事前知識がない場合、Lepskiの手法を用いて未知のスケールに適応可能であるが、追加の仮定がなければ観測可能な信頼区間は得られなくなる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。