Skip to main content
QUICK REVIEW

[论文解读] Generative Data Assimilation of Sparse Weather Station Observations at Kilometer Scales

Peter Manshausen, Yair Cohen|arXiv (Cornell University)|Jun 19, 2024
Meteorological Phenomena and Simulations被引用 4
一句话总结

该论文提出了一种基于分数的扩散模型,用于在3公里分辨率下对稀疏气象站观测数据进行生成式数据同化,实现了快速、可扩展且物理上合理的地表风场与降水场重建。在留出的气象站上,其均方根误差(RMSE)比业务运行的HRRR系统低10%,证明了该方法在无需微调的情况下实现低延迟公里尺度集合再分析的可行性。

ABSTRACT

Data assimilation of observational data into full atmospheric states is essential for weather forecast model initialization. Recently, methods for deep generative data assimilation have been proposed which allow for using new input data without retraining the model. They could also dramatically accelerate the costly data assimilation process used in operational regional weather models. Here, in a central US testbed, we demonstrate the viability of score-based data assimilation in the context of realistically complex km-scale weather. We train an unconditional diffusion model to generate snapshots of a state-of-the-art km-scale analysis product, the High Resolution Rapid Refresh. Then, using score-based data assimilation to incorporate sparse weather station data, the model produces maps of precipitation and surface winds. The generated fields display physically plausible structures, such as gust fronts, and sensitivity tests confirm learnt physics through multivariate relationships. Preliminary skill analysis shows the approach already outperforms a naive baseline of the High-Resolution Rapid Refresh system itself. By incorporating observations from 40 weather stations, 10% lower RMSEs on left-out stations are attained. Despite some lingering imperfections such as insufficiently disperse ensemble DA estimates, we find the results overall an encouraging proof of concept, and the first at km-scale. It is a ripe time to explore extensions that combine increasingly ambitious regional state generators with an increasing set of in situ, ground-based, and satellite remote sensing data streams.

研究动机与目标

  • 开发一种可扩展、低延迟的方法,利用稀疏气象站数据初始化高分辨率大气状态。
  • 评估基于分数的扩散模型是否能够从公里尺度再分析数据中学习到物理上合理的大气动力学。
  • 证明该模型可在无需微调的情况下适应新观测,从而实现对多样化数据流的灵活同化。
  • 评估生成场的性能与HRRR等业务基准的对比,特别是在降水和地表风估计方面的表现。
  • 探索生成模型作为复杂、计算成本高昂的数据同化流程替代品的潜力。

提出的方法

  • 在High Resolution Rapid Refresh(HRRR)再分析数据集的快照上预训练扩散模型,以生成3公里分辨率的地表场。
  • 应用基于分数的数据同化(SDA)框架,利用去噪分数函数将扩散模型与稀疏气象站观测条件化。
  • SDA框架使用噪声调度和去噪网络,通过迭代方式逐步优化生成场,使其与观测数据保持一致。
  • 模型通过最小化损失函数进行训练,使预测分数与噪声数据分布的真实分数对齐。
  • 通过对角协方差矩阵 $\sqrt{\Sigma_y}$ 建模观测不确定性,其数值通过经验调优确定。
  • 该方法支持不同去噪步数和校正迭代次数的推理,从而在准确性和速度之间实现权衡。
Figure 1: Denoiser training and data assimilation with SDA. a) During the training of the denoiser, noise is added to the training data at different levels, parameterized by time $t\in[0,1]$ . The training objective for the denoiser $D$ is to reconstruct the training data, given the noisy state and
Figure 1: Denoiser training and data assimilation with SDA. a) During the training of the denoiser, noise is added to the training data at different levels, parameterized by time $t\in[0,1]$ . The training objective for the denoiser $D$ is to reconstruct the training data, given the noisy state and

实验结果

研究问题

  • RQ1在稀疏观测条件下,基于再分析数据训练的扩散模型能否生成物理上合理的3公里尺度地表天气场?
  • RQ2该模型是否能学习到与大气物理一致的多变量关系,例如风场与降水场之间的关系?
  • RQ3在未见气象站上,该模型是否能在不微调的情况下实现优于业务HRRR系统的精度?
  • RQ4模型性能如何随噪声调度和校正步数等超参数变化?
  • RQ5该框架能否扩展以整合多样化观测流,包括地面观测和卫星数据?

主要发现

  • 该模型成功生成了3公里分辨率下物理上合理的地表风场与降水场,包括阵风锋等特征。
  • 敏感性测试表明,该模型准确捕捉了与大气动力学一致的多变量关系。
  • 在留出的气象站上,该模型的RMSE比业务HRRR系统低10%,表明其性能更优。
  • 该方法对新观测具有鲁棒性,无需微调,可实现对新数据流的快速适应。
  • 尽管集合展度存在一些局限,但结果首次成功验证了公里尺度生成式数据同化的可行性。
  • 超参数调优显示,较低的 $\Gamma$(0.001)和较高的去噪步数可提升性能,尤其在降水预测方面。
Figure 2: Assimilating increasingly sparse and noisy data. Columns show the different variables 10u, 10v, and tp for different study cases. In row one, we show HRRR data of 2017-05-28 03:00 UTC. Rows two and three show this data subsampled to 1.6% and 0.3%, respectively, in a regular grid (shown as
Figure 2: Assimilating increasingly sparse and noisy data. Columns show the different variables 10u, 10v, and tp for different study cases. In row one, we show HRRR data of 2017-05-28 03:00 UTC. Rows two and three show this data subsampled to 1.6% and 0.3%, respectively, in a regular grid (shown as

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。