[Paper Review] Testing whether a Learning Procedure is Calibrated
This paper proposes a general framework for testing whether a learning procedure—such as Bayesian or robust Bayesian inference—is calibrated, meaning the true parameters are plausible draws from its output distribution. Using simulation-based hypothesis testing with the Kolmogorov–Smirnov test on predictive integral transforms, the method assesses strong and weak calibration, demonstrating that fractional posteriors can improve calibration under model misspecification.
A learning procedure takes as input a dataset and performs inference for the parameters $θ$ of a model that is assumed to have given rise to the dataset. Here we consider learning procedures whose output is a probability distribution, representing uncertainty about $θ$ after seeing the dataset. Bayesian inference is a prime example of such a procedure, but one can also construct other learning procedures that return distributional output. This paper studies conditions for a learning procedure to be considered calibrated, in the sense that the true data-generating parameters are plausible as samples from its distributional output. A learning procedure whose inferences and predictions are systematically over- or under-confident will fail to be calibrated. On the other hand, a learning procedure that is calibrated need not be statistically efficient. A hypothesis-testing framework is developed in order to assess, using simulation, whether a learning procedure is calibrated. Several vignettes are presented to illustrate different aspects of the framework.
Motivation & Objective
- To define and formalize the concept of calibration for learning procedures that output probability distributions over parameters.
- To address the lack of a general, mathematically precise definition of calibration applicable to arbitrary learning procedures beyond Bayesian inference.
- To develop a practical, simulation-based hypothesis testing framework to assess whether a learning procedure is calibrated.
- To distinguish between strong and weak calibration, enabling both rigorous evaluation and more accessible testing.
- To illustrate the framework across diverse learning procedures, including Bayesian, variational, and robust Bayesian methods, under model misspecification.
Proposed method
- Formalize a learning procedure as a measurable map from data to a posterior distribution over parameters, under mild regularity conditions.
- Define strong calibration as the condition that the predictive integral transform (PIT) of the true parameter under the posterior is uniformly distributed on [0,1].
- Define weak calibration as the condition that the PIT distribution is stochastically equal to the uniform distribution under repeated sampling.
- Use the Kolmogorov–Smirnov (KS) test to evaluate the null hypothesis that the PITs are uniformly distributed, enabling statistical testing of calibration.
- Apply the framework to simulated data under known data-generating models, including misspecified likelihoods, to test calibration of various learning procedures.
- Use Monte Carlo sampling to estimate the PIT distribution and compute test statistics for calibration assessment.
Experimental results
Research questions
- RQ1What constitutes a general, mathematically rigorous definition of calibration for learning procedures that output probability distributions?
- RQ2How can one test whether a learning procedure is calibrated in practice, especially when the posterior is intractable or the model is misspecified?
- RQ3To what extent do alternative learning procedures—such as fractional posteriors or variational Bayes—improve calibration compared to standard Bayesian inference?
- RQ4Can the proposed framework detect overconfidence or miscalibration in learning procedures under model misspecification?
- RQ5How do strong and weak calibration differ in their practical implications and testability?
Key findings
- Fractional posterior methods with exponent $ t = 0.1, 0.2, 0.3 $ show improved calibration over standard Bayesian inference under likelihood misspecification, as evidenced by more uniform PIT distributions.
- The KS test on PITs successfully rejects the null hypothesis of strong calibration for standard Bayesian inference under misspecification, indicating miscalibration.
- For fractional posteriors, the PITs are more uniformly distributed, and the KS test p-values are higher, suggesting better calibration, especially for $ t = 0.1 $.
- The framework detects overconfidence in standard Bayesian inference through significant deviations of PITs from uniformity, even when posterior predictive checks may appear acceptable.
- Weak calibration is more easily testable than strong calibration, and the framework provides a practical pathway to assess calibration without requiring exact posterior computation.
- The results suggest that calibration is a distinct and critical desideratum that should be evaluated separately from statistical efficiency or predictive performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.