[论文解读] Regression Diagnostics meets Forecast Evaluation: Conditional Calibration, Reliability Diagrams, and Coefficient of Determination
本文通过将回归诊断与预测评估相联系,提出了一套统一的框架,用于评估回归与预测中的校准性。该框架提出T-可靠性图、将评分分解为偏差(MCB)、区分度(DSC)和不确定性(UNC)三部分,以及一个适用于任意可识别函数T的通用决定系数R∗ ∈ [0,1],其估计通过非参数保序回归与池相邻违反者(PAV)算法实现,具有稳定性。
Model diagnostics and forecast evaluation are two sides of the same coin. A common principle is that fitted or predicted distributions ought to be calibrated or reliable, ideally in the sense of auto-calibration, where the outcome is a random draw from the posited distribution. For binary responses, this is the universal concept of reliability. For real-valued outcomes, a general theory of calibration has been elusive, despite a recent surge of interest in distributional regression and machine learning. We develop a framework rooted in probability theory, which gives rise to hierarchies of calibration, and applies to both predictive distributions and stand-alone point forecasts. In a nutshell, a prediction - distributional or single-valued - is conditionally T-calibrated if it can be taken at face value in terms of the functional T. Whenever T is defined via an identification function - as in the cases of threshold (non) exceedance probabilities, quantiles, expectiles, and moments - auto-calibration implies T-calibration. We introduce population versions of T-reliability diagrams and revisit a score decomposition into measures of miscalibration (MCB), discrimination (DSC), and uncertainty (UNC). In empirical settings, stable and efficient estimators of T-reliability diagrams and score components arise via nonparametric isotonic regression and the pool-adjacent-violators algorithm. For in-sample model diagnostics, we propose a universal coefficient of determination, $$ ext{R}^\ast = \frac{ ext{DSC}- ext{MCB}}{ ext{UNC}},$$ that nests and reinterprets the classical $ ext{R}^2$ in least squares (mean) regression and its natural analogue $ ext{R}^1$ in quantile regression, yet applies to T-regression in general, with MCB $\geq 0$, DSC $\geq 0$, and $ ext{R}^\ast \in [0,1]$ under modest conditions.
研究动机与目标
- 通过校准的统一原则,统一回归诊断与预测评估。
- 为预测分布与点预测的条件校准建立严谨的理论框架。
- 将决定系数R²推广至任意可识别函数T(包括均值、分位数与期望值),实现广义化。
- 通过非参数保序回归,提供实用、稳定且可复现的校准性评估工具。
- 通过将评分分解为MCB、DSC与UNC,实现对各类预测类型的预测质量一致评估。
提出的方法
- 基于概率论与函数依赖关系,提出校准概念的分层体系——自校准、强阈值校准、条件校准与边际校准。
- 将T-校准定义为:在给定预测的条件下,预测分布的函数T的期望值等于实际观测结果。
- 通过非参数保序回归,将观测到的T值对预测的T值进行回归,构建总体与经验T-可靠性图。
- 应用池相邻违反者(PAV)算法估计校准后的T值,确保单调性与稳定性。
- 推导出评分分解公式:平均评分 = MCB − DSC + UNC,其中MCB ≥ 0,DSC ≥ 0,UNC > 0。
- 提出通用决定系数R∗ = (DSC − MCB) / UNC,其在温和条件下位于[0,1]区间,广义化了R²与R₁。
实验结果
研究问题
- RQ1如何为不同函数T(如均值、分位数或期望值)的预测分布与点预测,一致地定义与评估校准性?
- RQ2自校准与T-校准之间的理论关系是什么?在何种条件下自校准可推出T-校准?
- RQ3在一般评分函数框架下,如何一致地分解偏差、区分度与不确定性?
- RQ4能否为任意可识别函数T定义一个广义化R²与R₁的通用决定系数R∗?
- RQ5在经验设定下,如何构建T-可靠性图与评分分量的可靠、稳定且可复现的估计器?
主要发现
- 所提出的决定系数R∗在温和正则性条件下有界于[0,1],并广义化了最小二乘回归中的经典R²与分位数回归中的R₁。
- T-可靠性图通过将观测T值对预测T值进行非参数保序回归构建,确保估计的稳定与可复现。
- 对于任意一致评分函数,评分分解式S̄ = MCB − DSC + UNC成立,其中MCB ≥ 0,DSC ≥ 0,UNC > 0。
- 对于均值函数,MCB分量对应于均方误差(MSE)减去区分度分量,而R∗捕捉了可解释变异的比例。
- 通过CORP(一致、最优、可复现、基于PAV)方法对T-可靠性图与评分分量进行经验估计,确保统计效率、可复现性与稳定性。
- 该框架可普遍适用于任意可识别函数T,包括二值事件概率、分位数、矩与期望值,且不依赖于预测生成方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。