Skip to main content
QUICK REVIEW

[Paper Review] Evaluating model calibration in classification

Juozas Vaicenavičius, David Widmann|arXiv (Cornell University)|Feb 19, 2019
Software Reliability and Analysis ResearchComputer Science90 citations
TL;DR

This paper develops a general theoretical framework for evaluating calibration in probabilistic classifiers and introduces refined methods to quantify and visualize miscalibration, including multidimensional reliability diagrams.

ABSTRACT

Probabilistic classifiers output a probability distribution on target classes rather than just a class prediction. Besides providing a clear separation of prediction and decision making, the main advantage of probabilistic models is their ability to represent uncertainty about predictions. In safety-critical applications, it is pivotal for a model to possess an adequate sense of uncertainty, which for probabilistic classifiers translates into outputting probability distributions that are consistent with the empirical frequencies observed from realized outcomes. A classifier with such a property is called calibrated. In this work, we develop a general theoretical calibration evaluation framework grounded in probability theory, and point out subtleties present in model calibration evaluation that lead to refined interpretations of existing evaluation techniques. Lastly, we propose new ways to quantify and visualize miscalibration in probabilistic classification, including novel multidimensional reliability diagrams.

Motivation & Objective

  • Motivate the importance of calibrated probability estimates in safety-critical classification tasks.
  • Develop a general probabilistic calibration evaluation framework grounded in probability theory.
  • Identify subtleties in existing calibration evaluation techniques that affect interpretation.
  • Propose new metrics and visualization tools to quantify and visualize miscalibration.

Proposed method

  • Formulate a probabilistic calibration evaluation framework based on probability theory.
  • Analyze subtleties in existing calibration metrics and evaluation procedures.
  • Introduce novel visualization techniques for miscalibration, including multidimensional reliability diagrams.

Experimental results

Research questions

  • RQ1How can calibration of probabilistic classifiers be rigorously defined and evaluated?
  • RQ2What subtleties do common calibration evaluation methods have, and how can they be refined?
  • RQ3What new metrics and visual tools can effectively quantify and illustrate miscalibration in multiclass settings?

Key findings

  • A theoretical framework for calibration evaluation grounded in probability theory is proposed.
  • Identification of subtleties in existing calibration evaluation approaches that lead to refined interpretations.
  • Introduction of new quantification and visualization methods for miscalibration, including multidimensional reliability diagrams.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.