Skip to main content
QUICK REVIEW

[论文解读] Statistics of extreme events in coarse-scale climate simulations via machine learning correction operators trained on nudged datasets

Alexis-Tzianni Charalampopoulos, Shixuan Zhang|arXiv (Cornell University)|Apr 4, 2023
Climate variability and models被引用 5
一句话总结

该论文提出了一种非侵入式机器学习校正框架,通过在弱耦合的粗分辨率E3SM模拟上训练神经算子,以校正自由运行气候模型的输出。通过利用ERA5再分析数据作为参考,该方法准确再现了极端事件统计特征——包括热带气旋频率和大气河流事件——相较于未经校正的粗分辨率模拟,显著提升了性能,同时保持了稳定性并具备对未见数据的泛化能力。

ABSTRACT

This work presents a systematic framework for improving the predictions of statistical quantities for turbulent systems, with a focus on correcting climate simulations obtained by coarse-scale models. While high resolution simulations or reanalysis data are available, they cannot be directly used as training datasets to machine learn a correction for the coarse-scale climate model outputs, since chaotic divergence, inherent in the climate dynamics, makes datasets from different resolutions incompatible. To overcome this fundamental limitation we employ coarse-resolution model simulations nudged towards high quality climate realizations, here in the form of ERA5 reanalysis data. The nudging term is sufficiently small to not pollute the coarse-scale dynamics over short time scales, but also sufficiently large to keep the coarse-scale simulations close to the ERA5 trajectory over larger time scales. The result is a compatible pair of the ERA5 trajectory and the weakly nudged coarse-resolution E3SM output that is used as input training data to machine learn a correction operator. Once training is complete, we perform free-running coarse-scale E3SM simulations without nudging and use those as input to the machine-learned correction operator to obtain high-quality (corrected) outputs. The model is applied to atmospheric climate data with the purpose of predicting global and local statistics of various quantities of a time-period of a decade. Using datasets that are not employed for training, we demonstrate that the produced datasets from the ML-corrected coarse E3SM model have statistical properties that closely resemble the observations. Furthermore, the corrected coarse-scale E3SM output for the frequency of occurrence of extreme events, such as tropical cyclones and atmospheric rivers are presented. We present thorough comparisons and discuss limitations of the approach.

研究动机与目标

  • 在不改变其动力方程的前提下,提升粗分辨率气候模拟的统计准确性。
  • 解决高分辨率与粗分辨率气候模拟之间因混沌发散而导致的直接训练机器学习模型时的数据不匹配问题。
  • 开发一种非侵入式、后处理校正方法,以保持自由运行粗分辨率模拟的稳定性和物理一致性。
  • 仅使用粗分辨率模型输出,实现对非高斯统计特性及极端事件频率(如热带气旋和大气河流)的准确预测。
  • 在10年真实大气数据上验证该方法,展示其与ERA5再分析基准的一致性。

提出的方法

  • 使用弱耦合的E3SM模拟训练机器学习算子,这些模拟被引导沿ERA5再分析轨迹运行,以生成兼容的输入-输出对。
  • 以耦合E3SM输出作为输入,ERA5数据作为目标输出,训练基于LSTM的神经网络(特定为算子)以学习校正映射关系。
  • 在后处理阶段应用训练好的校正算子于自由运行的E3SM模拟(无耦合)中,以生成高保真度的统计输出。
  • 对神经网络输出应用谱校正,进一步使能量谱与参考数据对齐。
  • 采用25个子区域划分方案,评估全球范围内的平均水汽输送(IVT)统计特性。
  • 使用训练中未使用的独立ERA5数据验证方法,重点关注风速、温度、湿度及极端事件的全球与局部统计特性。

实验结果

研究问题

  • RQ1基于耦合粗分辨率模拟训练的机器学习校正算子,能否准确再现风速、温度和湿度等大气变量的非高斯统计特性?
  • RQ2该方法在多大程度上可校正粗分辨率气候模型中对热带气旋频率的低估问题?
  • RQ3对于大气河流频率和强度等极端事件统计,ML校正输出与ERA5再分析数据的匹配程度如何?
  • RQ4非侵入式、后处理校正方法是否能保持稳定性,避免在线校正方案中常见的数值不稳定性?
  • RQ5该方法是否能泛化至无耦合的自由运行模拟,同时仍能生成准确的长期统计特性?

主要发现

  • ML校正后的E3SM模拟在平均水汽输送(IVT)预测上的均方误差相比耦合数据集降低了25%,且与ERA5数据的吻合度更优。
  • 校正模型在2007–2017年间预测全球共411个热带气旋,接近ERA5参考值488个,而未经校正的CLIM模拟仅预测404个。
  • 该模型成功捕捉了风速分量(U、V)、温度(T)和比湿(Q)的概率密度函数的非高斯尾部特征,优于CLIM和耦合模拟。
  • 校正算子准确再现了大气河流的空间分布与频率,尤其在北太平洋和北大西洋等关键区域表现突出。
  • 即使在粗分辨率模型(CLIM)因无法生成足够涡度而无法在大西洋形成热带气旋时,该方法仍能稳健预测极端事件统计。
  • 谱校正显著改善了校正输出的能量谱与参考ERA5数据的一致性,提升了整体统计保真度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。