Skip to main content
QUICK REVIEW

[论文解读] Forecast score distributions with imperfect observations

Julie Bessac, Philippe Naveau|arXiv (Cornell University)|Jun 10, 2018
Meteorological Phenomena and Simulations参考文献 18被引用 4
一句话总结

该论文提出了一种新颖的评分规则框架,通过将观测误差建模为隐变量和条件期望,考虑了验证数据中的不确定性,从而在观测不完美时提升了预测评估效果。其核心贡献在于:使用完整的评分分布——尤其是通过Wasserstein距离——可显著增强对均值评分的区分能力,尤其在分布偏斜或重尾以及数据噪声较大的情况下。

ABSTRACT

The classical paradigm of scoring rules is to discriminate between two different forecasts by comparing them with observations. The probability distribution of the observed record is assumed to be perfect as a verification benchmark. In practice, however, observations are almost always tainted by errors and uncertainties. If the yardstick used to compare forecasts is imprecise, one can wonder whether such types of errors may or may not have a strong influence on decisions based on classical scoring rules. We propose a new scoring rule scheme in the context of models that incorporate errors of the verification data. We rely on existing scoring rules and incorporate uncertainty and error of the verification data through a hidden variable and the conditional expectation of scores when they are viewed as a random variable. The proposed scoring framework is compared to scores used in practice, and is expressed in various setups, mainly an additive Gaussian noise model and a multiplicative Gamma noise model. By considering scores as random variables one can access the entire range of their distribution. In particular we illustrate that the commonly used mean score can be a misleading representative of the distribution when this latter is highly skewed or have heavy tails. In a simulation study, through the power of a statistical test and the computation of Wasserstein distances between scores distributions, we demonstrate the ability of the newly proposed score to better discriminate between forecasts when verification data are subject to uncertainty compared with the scores used in practice. Finally, we illustrate the benefit of accounting for the uncertainty of the verification data into the scoring procedure on a dataset of surface wind speed from measurements and numerical model outputs.

研究动机与目标

  • 解决预测评估中的关键空白:验证数据常被假设为无误差,但现实中存在仪器、再分析或间接测量带来的不确定性。
  • 开发一种统计上严谨的方法,将验证数据的不确定性纳入评分规则,而无需事先了解真实过程的精确信息。
  • 在观测不完美时,提升预测评分的区分能力,尤其当均值评分因偏度或重尾而失效时。
  • 通过模拟和真实风速数据验证,表明完整评分分布(特别是Wasserstein距离)在数据不确定性下,优于基于均值的评分,能更有效地检测预测差异。

提出的方法

  • 该框架将真实(隐藏)状态视为潜变量,并将观测数据视为围绕其条件分布,引入如加性高斯噪声和乘性伽马噪声等误差模型。
  • 将评分视为随机变量,并推导其在给定观测数据下的条件期望,从而实现从验证数据到评分函数的不确定性传播。
  • 使用广义逆(分位函数)计算评分分布之间的1-Wasserstein距离,量化完整评分行为的差异。
  • 为保证可计算性,该方法假设已知误差模型(如正态或伽马分布),并在这些假设下推导出条件评分分布的显式表达式。
  • 通过具有已知误差结构的合成数据验证该框架,并将其应用于来自ASOS和WRF模型输出的真实地表风速数据。
  • 采用统计功效分析和Wasserstein距离比较,评估不同评分方案下的区分性能。

实验结果

研究问题

  • RQ1验证数据中的误差如何影响经典评分规则的可靠性(这些规则假设观测完美无误)?
  • RQ2能否构建一种评分规则框架,通过隐变量和条件期望显式考虑验证数据的不确定性?
  • RQ3与仅使用均值相比,使用完整的评分分布是否能提升在验证数据存在噪声时区分不同预测的能力?
  • RQ4在数据不确定性下,评分分布之间的Wasserstein距离与均值评分相比,在检测预测差异方面表现如何?
  • RQ5所提出的校正评分对观测误差是否具有鲁棒性,尤其当真实预测值接近观测值时?

主要发现

  • 当评分分布呈现偏斜或重尾时,均值评分无法有效代表评分分布,导致预测比较中产生误导性结论。
  • 使用完整评分分布,特别是通过1-Wasserstein距离,显著提升了评分规则的区分能力,优于仅依赖均值评分。
  • 评分分布之间的Wasserstein距离表现出更陡峭的梯度和更小的模糊区域,表明其在不完美数据条件下对预测差异具有更高的敏感性。
  • 所提出的校正评分(考虑验证数据不确定性)即使在验证数据存在噪声时,其Wasserstein距离也能接近真实最小值。
  • 在模拟实验中,新框架在检测预测差异方面表现出优于实际中常用标准评分的统计功效。
  • 该方法在真实地表风速数据上成功提升了预测评估效果,这些数据中存在仪器和再分析产品带来的观测误差。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。