Skip to main content
QUICK REVIEW

[论文解读] Weighted scoring rules and hypothesis testing

Hajo Holzmann, Bernhard Klar|arXiv (Cornell University)|Jan 1, 2016
Advanced Statistical Methods and Models参考文献 11被引用 3
一句话总结

本文提出了一种通用框架,用于基于条件密度构建严格局部适当的加权评分规则,从而实现对感兴趣区域的聚焦预测评估。研究证明,Diks 等人(2011)提出的右删失似然规则在独立同分布设定下可产生最优假设检验,表明加权评分规则能够可靠地识别出在特定区域内表现更优的预测,即使其在该区域外表现不佳。

ABSTRACT

We discuss weighted scoring rules for forecast evaluation and their connection to hypothesis testing. First, a general construction principle for strictly locally proper weighted scoring rules based on conditional densities and scoring rules for probability forecasts is proposed. We show how likelihood-based weighted scoring rules from the literature fit into this framework, and also introduce a weighted version of the Hyvärinen score, which is a local scoring rule in the sense that it only depends on the forecast density and its derivatives at the observation, and does not require evaluation of integrals. Further, we discuss the relation to hypothesis testing. Using a weighted scoring rule introduces a censoring mechanism, in which the form of the density is irrelevant outside the region of interest. For the resulting testing problem with composite null - and alternative hypotheses, we construct optimal tests, and identify the associated weighted scoring rule. As a practical consequence, using a weighted scoring rule allows to decide in favor of a forecast which is superior to a competing forecast on a region of interest, even though it may be inferior outside this region. A simulation study and an application to financial time-series data illustrate these findings.

研究动机与目标

  • 开发一种严格局部适当的加权评分规则的通用构建原则,以强调在特定感兴趣区域内的性能。
  • 通过确保理论上的适当性,解决加权评分规则偏向于在感兴趣区域内质量更高的预测的问题。
  • 将加权评分规则与复合原假设和备择假设下的最优假设检验相联系。
  • 证明即使在预测表现不佳的区域外,加权评分规则仍能可靠识别出在感兴趣区域内表现更优的预测。
  • 通过模拟和金融时间序列数据,评估加权评分规则的实际表现。

提出的方法

  • 提出一种利用条件密度和概率预测的适当评分规则来构建加权评分规则的通用方法。
  • 引入一种加权版本的 Hyvärinen 评分,该评分具有局部性,通过依赖观测点处的概率密度导数来避免积分计算。
  • 证明 Diks 等人(2011)和 Pelenis(2014)提出的基于似然的加权评分规则均符合所提出的框架。
  • 将预测比较重新表述为一个假设检验问题,其中原假设和备择假设为复合形式,而感兴趣区域充当删失机制。
  • 基于删失似然规则构建最优单侧检验,证明其在独立同分布设定下对应于最优检验。
  • 在简化框架中使用 Diebold-Mariano 检验,比较加权评分规则下的预测准确性。

实验结果

研究问题

  • RQ1如何构建加权评分规则,以确保其适当性,同时将评估重点集中于特定感兴趣区域?
  • RQ2加权评分规则与预测评估中最优假设检验之间有何关联?
  • RQ3即使预测在感兴趣区域外表现不佳,加权评分规则是否能可靠识别出在该区域内表现更优的预测?
  • RQ4与标准评分规则相比,不同加权评分规则(如删失似然规则和惩罚似然评分)在实际中的表现如何?
  • RQ5在仅关注区域性能的设定下,使用加权评分规则的理论依据是什么?

主要发现

  • Diks 等人(2011)提出的删失似然规则在独立同分布设定下可产生最优单侧检验,确立了其在区域特定预测评估中的理论最优性。
  • Pelenis(2014)提出的惩罚似然评分具有偏好保持性和适当性,在模拟和真实数据中表现良好,支持其实际应用价值。
  • 在德意志银行收益率分析中,偏斜- t 分布 GARCH 模型在预测损失方面显著优于 t-GARCH 模型(p 值 < 0.05),而 t-GARCH 模型在预测收益方面表现更优。
  • 模拟研究证实,即使全局性能在感兴趣区域外相似或更差,加权评分规则仍能检测到区域内的预测优越性。
  • 残差的可视化检查显示,偏斜-t GARCH 模型对左尾的拟合优于 t-GARCH 模型,而 t-GARCH 模型对右尾的拟合更优,与检验结果一致。
  • 尽管适当的加权评分规则在理论上具有吸引力,但模拟结果表明,当比较模型设定错误的情况时,其相对于非加权规则并无系统性改进——凸显了构建区域特定评估框架的必要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。