[Paper Review] Transferability and explainability of deep learning emulators for regional climate model projections: Perspectives for future applications
This study evaluates the transferability and explainability of deep learning emulators for regional climate model (RCM) projections, comparing two approaches—Perfect Prognosis (PP) and Model Output Statistics (MOS)—using eXplainable AI (XAI) techniques. It finds that PP emulators exhibit better physical consistency and soft transferability across GCMs, while MOS emulators are GCM-dependent and lack physical coherence, limiting their generalization to new climate models.
Regional climate models (RCMs) are essential tools for simulating and studying regional climate variability and change. However, their high computational cost limits the production of comprehensive ensembles of regional climate projections covering multiple scenarios and driving Global Climate Models (GCMs) across regions. RCM emulators based on deep learning models have recently been introduced as a cost-effective and promising alternative that requires only short RCM simulations to train the models. Therefore, evaluating their transferability to different periods, scenarios, and GCMs becomes a pivotal and complex task in which the inherent biases of both GCMs and RCMs play a significant role. Here we focus on this problem by considering the two different emulation approaches proposed in the literature (PP and MOS, following the terminology introduced in this paper). In addition to standard evaluation techniques, we expand the analysis with methods from the field of eXplainable Artificial Intelligence (XAI), to assess the physical consistency of the empirical links learnt by the models. We find that both approaches are able to emulate certain climatological properties of RCMs for different periods and scenarios (soft transferability), but the consistency of the emulation functions differ between approaches. Whereas PP learns robust and physically meaningful patterns, MOS results are GCM-dependent and lack physical consistency in some cases. Both approaches face problems when transferring the emulation function to other GCMs, due to the existence of GCM-dependent biases (hard transferability). This limits their applicability to build ensembles of regional climate projections. We conclude by giving some prospects for future applications.
Motivation & Objective
- To assess the transferability of deep learning emulators for regional climate model (RCM) projections across different Global Climate Models (GCMs), scenarios, and time periods.
- To evaluate the physical consistency of learned emulation functions using eXplainable AI (XAI) techniques.
- To compare the performance and reliability of two RCM emulation approaches: Perfect Prognosis (PP) and Model Output Statistics (MOS).
- To identify limitations in emulator generalization due to GCM-RCM bias mismatches and propose pathways for future application.
Proposed method
- Trained deep learning emulators on short RCM simulations driven by a single GCM to learn the mapping between large-scale atmospheric predictors and high-resolution surface variables (e.g., temperature, precipitation).
- Applied two distinct emulation frameworks: PP, which uses upscaled RCM fields as predictors, and MOS, which uses GCM fields as predictors.
- Employed eXplainable AI (XAI) methods such as saliency maps and feature attribution to interpret and validate the physical plausibility of learned relationships.
- Evaluated soft transferability by testing emulators on unseen GCMs, scenarios, and time periods.
- Assessed model performance using standard metrics and physical consistency of predictor patterns.
- Explored the impact of structural differences between GCMs and RCMs—such as aerosol representation and atmospheric physics—on emulator transferability.

Experimental results
Research questions
- RQ1Can deep learning emulators trained on one GCM-RCM combination reliably transfer to other GCMs and scenarios?
- RQ2How physically consistent are the relationships learned by PP and MOS emulators, as revealed by XAI techniques?
- RQ3Why do both PP and MOS emulators fail in hard transferability to new GCMs, and what role do GCM-RCM bias mismatches play?
- RQ4To what extent does the choice of training GCM affect the performance and interpretability of MOS-based emulators?
- RQ5What are viable strategies to improve the generalization of RCM emulators across diverse GCM-RCM combinations?
Key findings
- The PP approach learns robust, physically meaningful predictor patterns that are consistent across different driving GCMs, enhancing confidence in the emulation process.
- The MOS approach produces GCM-dependent patterns that lack physical consistency in some cases, making them less reliable for cross-GCM transfer.
- Both emulators exhibit soft transferability—performing reasonably well on new scenarios and periods—due to shared climatological features.
- Hard transferability fails when emulating new GCMs because of distinct RCM responses to GCM biases, especially in structural differences like aerosol representation.
- The mismatch in RCM response to different GCMs prevents reliable extrapolation, limiting the use of current emulators to fill the full GCM-RCM combination matrix.
- Future improvements may require training on diverse GCM biases or using reanalysis-based perfect boundary conditions, though both approaches involve trade-offs in accuracy and climate extrapolation.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.