[Paper Review] Climate-Invariant Machine Learning
The paper introduces climate-invariant ML, a physically informed framework that transforms inputs/outputs to keep learned mappings stable across climates, improving generalization of subgrid closures in three atmospheric models.
Projecting climate change is a generalization problem: we extrapolate the recent past using physical models across past, present, and future climates. Current climate models require representations of processes that occur at scales smaller than model grid size, which have been the main source of model projection uncertainty. Recent machine learning (ML) algorithms hold promise to improve such process representations, but tend to extrapolate poorly to climate regimes they were not trained on. To get the best of the physical and statistical worlds, we propose a new framework - termed "climate-invariant" ML - incorporating knowledge of climate processes into ML algorithms, and show that it can maintain high offline accuracy across a wide range of climate conditions and configurations in three distinct atmospheric models. Our results suggest that explicitly incorporating physical knowledge into data-driven models of Earth system processes can improve their consistency, data efficiency, and generalizability across climate regimes.
Motivation & Objective
- Motivate the need for ML models that generalize across climate regimes and reduce extrapolation errors in climate projections.
- Define and implement climate-invariant transformations to align input/output distributions across climates.
- Demonstrate robustness and generalization of climate-invariant ML closures across multiple atmospheric models and configurations.
- Explore how combining physical transformations with regularization improves data efficiency and transfer across climates.
Proposed method
- Introduce the climate-invariant mapping concept by transforming inputs/outputs so their distributions vary minimally across climates.
- Develop physically-informed transformations for inputs such as relative humidity, moist static energy buoyancy, and near-surface latent heat flux.
- Train and evaluate ML closures across three storm-resolving climate models (SPCAM3, SPCESM2, SAM) with cold, reference, and warm climate runs.
- Compare climate-invariant models (CI) to raw-data models (RD) and assess the impact of regularization (batch normalization, dropout).
- Use Earth-system relevant outputs including subgrid heating and moistening tendencies to assess generalization.
- Apply SHAP-based explainable AI to interpret why climate-invariant mappings generalize better across climates.
Experimental results
Research questions
- RQ1Can physically-informed input transformations stabilize ML mappings across climates for subgrid closures?
- RQ2Do climate-invariant transformations enable better out-of-distribution generalization across different climate regimes and configurations?
- RQ3How do regularization techniques interact with climate-invariant inputs to affect generalization?
- RQ4What structural differences do climate-invariant mappings exhibit compared to raw-data mappings?
Key findings
- Transforming inputs with climate-informed functions significantly improves cross-climate generalization of ML closures.
- Climate-invariant NNs trained on a cold climate generalize to warmer climates with substantially lower error than raw-data NNs.
- Combining climate-invariant transformations with regularization (BN, DP) yields the best generalization across climates and configurations.
- CI mappings tend to be more spatially local, aiding interpretability via SHAP analysis.
- CI models remain competitive with warm-climate training performance while avoiding extrapolation hazards typical of RD models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.