Skip to main content
QUICK REVIEW

[論文レビュー] Metrics of calibration for probabilistic predictions

Imanol Arrieta-Ibarra, Paman Gujral|arXiv (Cornell University)|May 19, 2022
Leaf Properties and Growth Measurement被引用数 8
ひとこと要約

本稿では、確率的予測のキャリブレーションを評価するための従来のビニングベースの指標よりも優れた代替手段として、経験的累積キャリブレーション誤差(ECCE)を提案する。観測確率と予測確率の累積差を分析することにより、ECCEは任意のビニングやカーネル帯域幅の選択を回避し、分解能とノイズのトレードオフなしに、より高い統計的信頼性とパワーを提供する。理論的保証が明確で、多様なデータセットにおいて優れた実験的性能を示す。

ABSTRACT

Predictions are often probabilities; e.g., a prediction could be for precipitation tomorrow, but with only a 30% chance. Given such probabilistic predictions together with the actual outcomes, "reliability diagrams" help detect and diagnose statistically significant discrepancies -- so-called "miscalibration" -- between the predictions and the outcomes. The canonical reliability diagrams histogram the observed and expected values of the predictions; replacing the hard histogram binning with soft kernel density estimation is another common practice. But, which widths of bins or kernels are best? Plots of the cumulative differences between the observed and expected values largely avoid this question, by displaying miscalibration directly as the slopes of secant lines for the graphs. Slope is easy to perceive with quantitative precision, even when the constant offsets of the secant lines are irrelevant; there is no need to bin or perform kernel density estimation. The existing standard metrics of miscalibration each summarize a reliability diagram as a single scalar statistic. The cumulative plots naturally lead to scalar metrics for the deviation of the graph of cumulative differences away from zero; good calibration corresponds to a horizontal, flat graph which deviates little from zero. The cumulative approach is currently unconventional, yet offers many favorable statistical properties, guaranteed via mathematical theory backed by rigorous proofs and illustrative numerical examples. In particular, metrics based on binning or kernel density estimation unavoidably must trade-off statistical confidence for the ability to resolve variations as a function of the predicted probability or vice versa. Widening the bins or kernels averages away random noise while giving up some resolving power. Narrowing the bins or kernels enhances resolving power while not averaging away as much noise.

研究の動機と目的

  • ECEなどのビニングベースのキャリブレーション指標が、任意のビン選択に依存するという、重大な限界を解消すること。これは結果を劇的に変える可能性があり、再現性を損なう。
  • ヒストグラムやカーネルベースのアプローチに内在する統計的信頼性と分解能のトレードオフを回避するキャリブレーション評価手法を開発すること。
  • 理論的根拠に基づき、パラメータフリーな方法として、誤差の測定が統計的に信頼性があり、実装が容易であるものを作り出すこと。
  • 理論的および実験的例を通じて、ECCEが有限標本サイズにおいても従来の指標を上回ることを示すこと。特に、誤差の検出能力に優れている。

提案手法

  • 予測値を順序付けた上で、観測結果と予測確率の差の累積和を計算することにより、累積プロットを構築する。
  • 2つの主要な指標を定義する:ECCE-MAD(ゼロからの最大絶対偏差)とECCE-R(偏差の範囲)。これらは累積プロットから直接的に誤差を測定する。
  • 漸近的理論を用いて、[17]の手法を応用し、再サンプリングやブートストラップに依存せずに、ECCEのP値を導出する。
  • パラメトリックでない、連続的な累積差関数を用いて、ヒストグラムベースのビニングやカーネル密度推定を置き換え、データの離散化を一切回避する。
  • 実世界のデータセット(例:ImageNet-1000、ブナザル、サングラス)および合成データに累積アプローチを適用し、さまざまなビニングスキーム下でのECEとの性能を比較する。
  • ECCEがビニング選択に対して不変であり、異なる標本サイズやデータ分布において一貫した性能を示すことを示す。

実験結果

リサーチクエスチョン

  • RQ1ビニングスキームの選択が、ECEなどの従来のキャリブレーション指標の信頼性と一貫性にどのように影響するか?
  • RQ2ビン化またはカーネルベースの手法に内在する分解能と統計的信頼性のトレードオフを回避できるキャリブレーション指標を開発できるか?
  • RQ3観測確率と予測確率の累積差を用いることで、どのような理論的および実験的利点が得られるか?
  • RQ4ECCEは、多様なデータセットや標本サイズにおいて、ECEに比べて誤差の検出に優れているか?
  • RQ5ECCEは、追加のチューニングや再サンプリングを必要とせず、漸近的に有効なP値を提供できるか?

主な発見

  • ECCEは、任意のビニングやカーネル帯域幅の選択を不要とし、それによりECEの値が異なる設定で一貫性のない結果を生じるのを回避する。
  • ECCE-MADとECCE-RはECEよりも統計的パワーが高く、分解能と信頼性のトレードオフがないため、同じレベルの誤差を検出するために必要な観測数が著しく少ない。
  • ImageNet-1000データセット(n = 1,281,167)において、ECCE-MADとECCE-Rはともに111.7 / σₙに達し、二重精度でP値がゼロとなった。これは、誤差の著しい統計的有意性を示している。
  • ブナザル(n = 1,300)およびサングラス(n = 1,300)データセットにおいて、ECCE-MADとECCE-Rはそれぞれ10.14および8.004 / σₙであった。P値がゼロとなったことから、誤差の強い統計的証拠が得られた。
  • ECCEアプローチはデータの順序に依存せず、チューニングパラメータも不要であるため、ECEに基づく手法よりもよりロバストで使いやすい。
  • 理論的分析により、ECCEがECEよりも少ない標本数で一致性とパワーを達成することが示された。特に大標本の極限において、ビニングに起因するバイアスと分散のトレードオフが存在しないためである。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。