Skip to main content
QUICK REVIEW

[Paper Review] Distillation and Interpretability of Ensemble Forecasts of ENSO Phase using Entropic Learning

Michael Groom, Davide Bassetti|arXiv (Cornell University)|Feb 15, 2026
Climate variability and models0 citations
TL;DR

The paper introduces a distillation framework that compresses a large ensemble of entropy-optimal Sparse Probabilistic Approximation (eSPA) models into compact, interpretable models for long-range ENSO phase forecasts up to 24 months, while preserving forecast skill.

ABSTRACT

This paper introduces a distillation framework for an ensemble of entropy-optimal Sparse Probabilistic Approximation (eSPA) models, trained exclusively on satellite-era observational and reanalysis data to predict ENSO phase up to 24 months in advance. While eSPA ensembles yield state-of-the-art forecast skill, they are harder to interpret than individual eSPA models. We show how to compress the ensemble into a compact set of "distilled" models by aggregating the structure of only those ensemble members that make correct predictions. This process yields a single, diagnostically tractable model for each forecast lead time that preserves forecast performance while also enabling diagnostics that are impractical to implement on the full ensemble. An analysis of the regime persistence of the distilled model "superclusters", as well as cross-lead clustering consistency, shows that the discretised system accurately captures the spatiotemporal dynamics of ENSO. By considering the effective dimension of the feature importance vectors, the complexity of the input space required for correct ENSO phase prediction is shown to peak when forecasts must cross the boreal spring predictability barrier. Spatial importance maps derived from the feature importance vectors are introduced to identify where predictive information resides in each field and are shown to include known physical precursors at certain lead times. Case studies of key events are also presented, showing how fields reconstructed from distilled model centroids trace the evolution from extratropical and inter-basin precursors to the mature ENSO state. Overall, the distillation framework enables a rigorous investigation of long-range ENSO predictability that complements real-time data-driven operational forecasts.

Motivation & Objective

  • Motivate interpretable long-range ENSO forecasting by addressing the opacity of large model ensembles.
  • Develop a distillation method to compress an eSPA ensemble into a single tractable model per lead time without losing predictive accuracy.
  • Demonstrate that distilled models retain predictive skill and reveal physically meaningful ENSO precursors through diagnostic maps and regime analysis.
  • Provide tools to investigate long-range ENSO predictability using a small set of interpretable models.

Proposed method

  • Train an ensemble of 50 eSPA classifiers for each lead time (1–24 months) using observational/reanalysis data.
  • Aggregate correct-prediction centroids from the ensemble into a dataset of superclusters after weighting by feature importance W^{(n)} and scaling.
  • Perform k-means clustering on the supercluster centroids to obtain K=12 superclusters per lead time.
  • Compute fuzzy affiliations Gamma^{(n)} via a convex quadratic program to map instances to superclusters.
  • Estimate conditional probabilities Lambda^{(n)} by solving a convex nonlinear program to relate supercluster affiliations to class probabilities.
  • Construct transition matrices P^{(n)} between superclusters using maximum likelihood on fuzzy affiliations, enabling Markov-chain analyses of ENSO dynamics.

Experimental results

Research questions

  • RQ1Can the ensemble of eSPA models be distilled into a compact set of interpretable models without sacrificing predictive skill for ENSO phase at various leads?
  • RQ2Do the distilled models reveal physically meaningful precursors and spatiotemporal dynamics consistent with ENSO mechanisms?
  • RQ3How do the input feature spaces and relation among clusters evolve across lead times, and what does this imply for predictability barriers?
  • RQ4Can diagnostic maps and derived transition structures (superclusters, Lambda, P) illuminate the pathways of ENSO development from precursors to mature states?

Key findings

  • Distilled models preserve probabilistic forecasting skill comparable to the full ensemble across lead times, while enabling interpretability.
  • Diagnostic maps from distilled models trace ENSO evolution from extratropical/inter-basin precursors to mature states.
  • Feature-importance-based analysis highlights predictors consistent with known ocean–atmosphere precursors at various lead times.
  • Transition matrices reveal dominant pathways between superclusters and quantify persistence and relaxation times (~6 months, aligning with Niño3.4 dynamics).
  • Case studies show reconstructed fields from distilled centroids reproduce the evolution of ENSO events and their precursors.
  • The approach provides a rigorous framework to study long-range ENSO predictability alongside real-time forecasts.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.