[Paper Review] Evaluating Probabilistic Classifiers: The Triptych
This paper introduces a triptych of diagnostic graphics—reliability diagrams, ROC curves, and Murphy diagrams—for evaluating probabilistic classifiers, enabling comprehensive assessment of calibration, discrimination, and overall predictive performance. The key contribution is a theoretically grounded framework that links these tools through score decomposition into miscalibration, discrimination, and uncertainty components, with equivalent rankings across metrics for calibrated forecasts.
Probability forecasts for binary outcomes, often referred to as probabilistic classifiers or confidence scores, are ubiquitous in science and society, and methods for evaluating and comparing them are in great demand. We propose and study a triptych of diagnostic graphics that focus on distinct and complementary aspects of forecast performance: The reliability diagram addresses calibration, the receiver operating characteristic (ROC) curve diagnoses discrimination ability, and the Murphy diagram visualizes overall predictive performance and value. A Murphy curve shows a forecast's mean elementary scores, including the widely used misclassification rate, and the area under a Murphy curve equals the mean Brier score. For a calibrated forecast, the reliability curve lies on the diagonal, and for competing calibrated forecasts, the ROC and Murphy curves share the same number of crossing points. We invoke the recently developed CORP (Consistent, Optimally binned, Reproducible, and Pool-Adjacent-Violators (PAV) algorithm based) approach to craft reliability diagrams and decompose a mean score into miscalibration (MCB), discrimination (DSC), and uncertainty (UNC) components. Plots of the DSC measure of discrimination ability versus the calibration metric MCB visualize classifier performance across multiple competitors. The proposed tools are illustrated in empirical examples from astrophysics, economics, and social science.
Motivation & Objective
- To address the lack of a unified, theoretically grounded framework for evaluating probabilistic classifiers across multiple performance dimensions.
- To resolve the challenge of choosing among competing evaluation metrics by proposing a coherent triptych of complementary diagnostic graphics.
- To provide a consistent, reproducible method for assessing calibration using the CORP approach based on PAV algorithms.
- To establish theoretical equivalence between sharpness, discrimination, and proper scoring rule performance under calibration.
- To enable practitioners to compare forecasts holistically using a combination of reliability, discrimination, and overall score performance.
Proposed method
- Proposes a triptych of diagnostic graphics: reliability diagrams (using CORP method for consistent, optimally binned, reproducible calibration assessment), ROC curves for discrimination, and Murphy diagrams for overall predictive performance.
- Employs the CORP framework to construct reliability diagrams via PAV-based binning, ensuring consistency and reproducibility in calibration evaluation.
- Uses score decomposition to break down the mean Brier score into three components: miscalibration (MCB), discrimination (DSC), and uncertainty (UNC).
- Represents proper scoring rules as Murphy curves, where the mean elementary score is plotted as a function of a threshold parameter θ, with area under the curve equal to the mean Brier score.
- Applies the concept of convex order to formalize sharpness, defining one forecast as sharper than another if it has higher expectation under all convex functions.
- Establishes theoretical equivalence between dominance in ROC curves, Murphy curves, and sharpness order under the assumption of calibration.
Experimental results
Research questions
- RQ1How can probabilistic classifiers be evaluated and compared across the three core dimensions of calibration, discrimination, and overall predictive performance?
- RQ2What is the theoretical relationship between miscalibration, discrimination, and uncertainty in probabilistic forecasting?
- RQ3Under what conditions do rankings of forecasts based on ROC curves, Murphy curves, and sharpness coincide?
- RQ4How can reliability diagrams be constructed in a consistent, reproducible, and optimally binned manner?
- RQ5What is the role of proper scoring rules in unifying the evaluation of probabilistic forecasts across different performance aspects?
Key findings
- For calibrated forecasts, the number of crossing points between ROC and Murphy curves is identical, providing a consistent basis for comparison.
- The area under a Murphy curve equals the mean Brier score, linking the integral of elementary scores to a widely used proper scoring rule.
- When forecasts are calibrated, a forecast that is sharper (in convex order) dominates all others in both ROC and Murphy curve comparisons.
- The difference between Murphy and ROC curve differences has the same number of sign changes, ensuring alignment in performance rankings.
- The CORP method produces reliability diagrams that are consistent, optimally binned, reproducible, and based on the PAV algorithm, improving upon traditional binning methods.
- Score decomposition into MCB, DSC, and UNC components enables detailed diagnostic analysis, with MCB and DSC plots offering a visual summary of classifier performance across multiple models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.