Skip to main content
QUICK REVIEW

[Paper Review] Weighted scoring rules and hypothesis testing

Hajo Holzmann, Bernhard Klar|arXiv (Cornell University)|Jan 1, 2016
Advanced Statistical Methods and Models11 references3 citations
TL;DR

This paper proposes a general framework for constructing strictly locally proper weighted scoring rules based on conditional densities, enabling focused forecast evaluation on regions of interest. It establishes that the censored likelihood rule from Diks et al. (2011) yields optimal hypothesis tests in i.i.d. settings, demonstrating that weighted scoring rules can reliably identify superior forecasts within specific domains—even when performance outside those regions is poor.

ABSTRACT

We discuss weighted scoring rules for forecast evaluation and their connection to hypothesis testing. First, a general construction principle for strictly locally proper weighted scoring rules based on conditional densities and scoring rules for probability forecasts is proposed. We show how likelihood-based weighted scoring rules from the literature fit into this framework, and also introduce a weighted version of the Hyvärinen score, which is a local scoring rule in the sense that it only depends on the forecast density and its derivatives at the observation, and does not require evaluation of integrals. Further, we discuss the relation to hypothesis testing. Using a weighted scoring rule introduces a censoring mechanism, in which the form of the density is irrelevant outside the region of interest. For the resulting testing problem with composite null - and alternative hypotheses, we construct optimal tests, and identify the associated weighted scoring rule. As a practical consequence, using a weighted scoring rule allows to decide in favor of a forecast which is superior to a competing forecast on a region of interest, even though it may be inferior outside this region. A simulation study and an application to financial time-series data illustrate these findings.

Motivation & Objective

  • To develop a general construction principle for strictly locally proper weighted scoring rules that emphasize performance in specific regions of interest.
  • To resolve the issue of improper weighted scoring rules favoring forecasts with higher mass in the region of interest, by ensuring theoretical propriety.
  • To connect weighted scoring rules to optimal hypothesis testing under composite null and alternative hypotheses.
  • To demonstrate that weighted scoring rules can identify superior forecasts within a region of interest, even if they underperform outside it.
  • To evaluate the practical performance of weighted scoring rules using simulations and financial time-series data.

Proposed method

  • Proposes a general construction of weighted scoring rules using conditional densities and proper scoring rules for probability forecasts.
  • Introduces a weighted version of the Hyvärinen score, which is local and avoids integral evaluation by relying on density derivatives at the observation.
  • Demonstrates that likelihood-based weighted scoring rules from Diks et al. (2011) and Pelenis (2014) fit within the proposed framework.
  • Reframes forecast comparison as a hypothesis testing problem with composite null and alternative hypotheses, where the region of interest acts as a censoring mechanism.
  • Constructs optimal one-sided tests based on the censored likelihood rule, showing it corresponds to the optimal test under the i.i.d. setting.
  • Uses the Diebold-Mariano test in a stylized framework to compare predictive accuracy under weighted scoring rules.

Experimental results

Research questions

  • RQ1How can weighted scoring rules be constructed to ensure propriety while focusing evaluation on a specific region of interest?
  • RQ2What is the connection between weighted scoring rules and optimal hypothesis testing in forecast evaluation?
  • RQ3Can weighted scoring rules reliably identify a forecast as superior within a region of interest, even if it underperforms outside that region?
  • RQ4How do different weighted scoring rules—such as the censored likelihood rule and penalized likelihood score—perform in practice compared to standard scoring rules?
  • RQ5What is the theoretical justification for using weighted scoring rules in settings where only regional performance matters?

Key findings

  • The censored likelihood rule from Diks et al. (2011) leads to optimal one-sided tests in the i.i.d. setting, establishing its theoretical optimality for region-specific forecast evaluation.
  • The penalized likelihood score by Pelenis (2014), which is preference-preserving and proper, performs well in simulations and real data, supporting its practical utility.
  • In the Deutsche Bank returns analysis, the skew-tdistribution GARCH model significantly outperforms the t-GARCH model for predicting losses (p-values < 0.05), while the t-GARCH model is superior for gains.
  • The simulation study confirms that weighted scoring rules can detect regional forecast superiority even when global performance is similar or worse outside the region of interest.
  • Visual inspection of residuals shows that the skew-t GARCH model fits the left tail better than the t-GARCH model, while the t-GARCH model fits the right tail better, aligning with the test results.
  • Despite the theoretical appeal of proper weighted scoring rules, simulations show no systematic improvement over non-weighted rules when comparing misspecified models—highlighting the need for region-specific evaluation frameworks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.