[Paper Review] A scoring framework for tiered warnings and multicategorical forecasts based on fixed risk measures
This paper introduces the FIxed Risk Multicategory (FIRM) Forecast Framework, a flexible family of scoring functions for evaluating ordered multicategorical forecasts—such as tiered weather warnings—based on fixed risk thresholds (α). The framework uses risk-consistent scoring matrices with user-defined weights and optional distance-based discounting for near misses, ensuring score optimization aligns with decision-theoretic directives, offering a robust alternative to traditional equitable scores that depend on variable thresholds and base rates.
The use of tiered warnings and multicategorical forecasts are ubiquitous in meteorological operations. Here, a flexible family of scoring functions is presented for evaluating the performance of ordered multicategorical forecasts. Each score has a risk parameter $\alpha$, selected for the specific use case, so that it is consistent with a forecast directive based on the fixed threshold probability $1-\alpha$ (equivalently, a fixed $\alpha$-quantile mapping). Each score also has use-case specific weights so that forecasters who accurately discriminate between categorical thresholds are rewarded in proportion to the weight for that threshold. A variation is presented where the penalty assigned to near misses or close false alarms is discounted, which again is consistent with directives based on fixed risk measures. The scores presented provide an alternative to many performance measures currently in use, whose optimal threshold probabilities for forecasting an event typically vary with each forecast case, and in the case of equitable scores are based around sample base rates rather than risk measures suitable for users.
Motivation & Objective
- To address the limitations of existing equitable scoring functions that optimize for variable threshold probabilities tied to sample base rates rather than user-specific risk tolerance.
- To develop a scoring framework that aligns forecast performance evaluation with fixed risk directives, such as issuing warnings when event probability exceeds 1−α.
- To provide a decision-theoretically consistent method for evaluating multicategorical forecasts, especially in public warning systems where false alarms and missed events carry asymmetric costs.
- To enable forecasters to be rewarded based on their ability to discriminate between high-impact thresholds, using category-specific weights.
- To incorporate distance-based discounting of penalties for near misses and close false alarms, reflecting real-world tolerance for forecast proximity.
Proposed method
- Proposes a family of consistent scoring functions parameterized by a risk level α, where forecasts are evaluated based on whether they contain the α-quantile of the predictive distribution.
- Uses a scoring matrix to assign penalties to forecast-observation pairs, with entries derived from elementary scoring functions for quantiles and expectiles.
- Incorporates user-specified weights (wi) to reward accurate discrimination at higher-priority thresholds.
- Introduces a discounting parameter 'a' to reduce penalties when observations fall within distance 'a' of the forecast category, modeling Huber quantiles or expectiles.
- Derives the framework from decision theory using the cost-loss model, ensuring that score optimization aligns with optimal user action under fixed risk measures.
- Applies the framework to real-world rainfall warning data, demonstrating parameter estimation via empirical calibration and signal detection theory.
Experimental results
Research questions
- RQ1Can a scoring function be constructed that remains consistent with a fixed risk threshold (1−α) across all forecast categories, rather than varying with base rates or sample frequencies?
- RQ2How can forecasters be incentivized to accurately discriminate between high-impact thresholds through category-specific weights in the scoring function?
- RQ3To what extent does discounting penalties for near misses or close false alarms improve the alignment of forecast scores with real-world user tolerance and decision-making?
- RQ4How can the risk parameter α be estimated from historical contingency tables or forecast-observation data when it is not explicitly defined in an existing warning system?
- RQ5Does varying the risk parameter α with lead time improve forecast system performance in multi-stage warning systems, and how should optimal thresholds be selected?
Key findings
- The FIRM scoring framework ensures that optimizing the score leads to forecast behavior consistent with a fixed risk directive, such as issuing a warning when the event probability exceeds 1−α.
- For well-calibrated forecast systems, the ratio of misses to false alarms is not a reliable indicator of performance, as it can be misleading when α is not aligned with the base rate.
- The signal detection theory-based estimator ˜α provides a more reliable estimate of α than the empirical estimator ˆα, especially at low base rates and with imperfectly calibrated systems.
- When α varies with threshold (e.g., α1=0.1, α2=0.9), forecasters may be incentivized to skip intermediate categories (e.g., C1), leading to suboptimal behavior and violating intuitive forecasting logic.
- In multi-lead-time warning systems, the optimal early warning threshold (β) is higher than the standard-time threshold (α), with β=0.65 for OCF and β=0.60 for Official in the NSW rainfall dataset, reflecting the higher cost of retracting warnings.
- The framework’s flexibility allows it to reduce the impact of warning fatigue and false alarm intolerance by discounting penalties for observations near forecast categories, particularly when a>0.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.