Skip to main content
QUICK REVIEW

[论文解读] Metrics of calibration for probabilistic predictions

Imanol Arrieta-Ibarra, Paman Gujral|arXiv (Cornell University)|May 19, 2022
Leaf Properties and Growth Measurement被引用 8
一句话总结

本文提出经验累积校准误差(ECCEs)作为传统基于分箱的校准度量指标的更优替代方案,用于评估概率预测的校准性。通过分析观测概率与预测概率之间的累积差异,ECCEs 避免了任意分箱或核带宽选择,提供更高的统计可靠性与检验效能,且无需在分辨率与噪声之间权衡,兼具严格的理论保证与在多样化数据集上的强劲实证表现。

ABSTRACT

Predictions are often probabilities; e.g., a prediction could be for precipitation tomorrow, but with only a 30% chance. Given such probabilistic predictions together with the actual outcomes, "reliability diagrams" help detect and diagnose statistically significant discrepancies -- so-called "miscalibration" -- between the predictions and the outcomes. The canonical reliability diagrams histogram the observed and expected values of the predictions; replacing the hard histogram binning with soft kernel density estimation is another common practice. But, which widths of bins or kernels are best? Plots of the cumulative differences between the observed and expected values largely avoid this question, by displaying miscalibration directly as the slopes of secant lines for the graphs. Slope is easy to perceive with quantitative precision, even when the constant offsets of the secant lines are irrelevant; there is no need to bin or perform kernel density estimation. The existing standard metrics of miscalibration each summarize a reliability diagram as a single scalar statistic. The cumulative plots naturally lead to scalar metrics for the deviation of the graph of cumulative differences away from zero; good calibration corresponds to a horizontal, flat graph which deviates little from zero. The cumulative approach is currently unconventional, yet offers many favorable statistical properties, guaranteed via mathematical theory backed by rigorous proofs and illustrative numerical examples. In particular, metrics based on binning or kernel density estimation unavoidably must trade-off statistical confidence for the ability to resolve variations as a function of the predicted probability or vice versa. Widening the bins or kernels averages away random noise while giving up some resolving power. Narrowing the bins or kernels enhances resolving power while not averaging away as much noise.

研究动机与目标

  • 解决基于分箱的校准度量指标(如 ECE)所依赖的任意分箱选择这一关键局限性,该选择可能导致结果发生剧烈变化并阻碍可复现性。
  • 开发一种校准评估方法,避免直方图或核密度方法中固有的统计置信度与分辨率之间的权衡。
  • 建立一种理论基础坚实、无需参数调节的校准度量方法,实现统计可靠性与实际可实施性的统一。
  • 通过理论分析与实证示例证明,ECCEs 在检测校准偏差方面优于传统指标,尤其在有限样本量下表现更优。

提出的方法

  • 通过计算有序预测中观测结果与预测概率之间差异的累积和,构建累积图。
  • 定义两个关键指标:ECCE-MAD(与零的绝对偏差最大值)和 ECCE-R(偏差范围),直接从累积图中量化校准偏差。
  • 利用渐近理论,通过 [17] 中的方法推导 ECCE 的 P 值,实现在无需重抽样或自助法的情况下进行统计显著性检验。
  • 以非参数化、连续的累积差异函数替代基于直方图的分箱与核密度估计,完整保留所有数据,避免离散化处理。
  • 将累积方法应用于真实世界数据集(如 ImageNet-1000、野猪、太阳镜)及合成数据,比较其在不同分箱方案下与 ECE 的性能表现。
  • 证明 ECCEs 对分箱选择具有不变性,并在不同样本量与数据分布下保持一致的性能表现。

实验结果

研究问题

  • RQ1分箱方案的选择如何影响传统校准度量指标(如 ECE)的可靠性与一致性?
  • RQ2能否开发一种校准度量指标,避免基于分箱或核密度方法中固有的分辨率与统计置信度之间的权衡?
  • RQ3使用观测概率与预测概率之间的累积差异进行校准评估,其理论与实证优势为何?
  • RQ4ECCEs 在检测不同数据集与样本量下的校准偏差方面,与 ECE 相比表现如何?
  • RQ5ECCEs 是否能提供渐近有效的 P 值以支持统计显著性检验,而无需额外调参或重抽样?

主要发现

  • ECCEs 消除了对任意分箱或核带宽选择的依赖,从而避免了因配置不同而导致 ECE 值不一致与不可靠的问题。
  • ECCE-MAD 与 ECCE-R 指标在统计效能上显著优于 ECE,由于不存在分辨率-置信度权衡,检测相同水平校准偏差所需样本量更少。
  • 在 ImageNet-1000 数据集(n = 1,281,167)中,ECCE-MAD 与 ECCE-R 均达到 111.7 / σₙ,渐近 P 值精确为零(双精度),表明校准偏差具有极强的统计显著性。
  • 在野猪(n = 1,300)与太阳镜(n = 1,300)数据集中,ECCE-MAD 与 ECCE-R 分别为 10.14 和 8.004 / σₙ,同样得到 P 值为零,证实了校准偏差存在强有力的统计证据。
  • ECCE 方法对数据排序具有不变性,且无需调参,使其比基于 ECE 的方法更具鲁棒性,也更易应用。
  • 理论分析证明,由于消除了分箱引起的偏差与方差权衡,ECCE 在样本量较大时的收敛一致性与检验效能均优于 ECE,尤其在大样本极限下表现更优。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。