Skip to main content
QUICK REVIEW

[Paper Review] Evaluating the Predictability of Selected Weather Extremes with Aurora, an AI Weather Forecast Model

Qin Huang, Moyan Liu|arXiv (Cornell University)|Mar 6, 2026
Tropical and Extratropical Cyclones Research0 citations
TL;DR

The paper evaluates Aurora, an AI weather forecast model, for predictability of selected weather extremes across tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation at lead times from 1 to 21 days, highlighting strong short-range skill but a subseasonal failure mode in extreme intensity.

ABSTRACT

AI weather foundation models now achieve forecast skill comparable to numerical weather prediction at far lower computational cost, yet their predictability for high-impact extremes across dynamical regimes remains uncertain. We evaluate Aurora using an event-based framework spanning tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation at lead times from 1 to 21 days. Aurora demonstrates strong short-range (1-7 day) skill across event types, including competitive tropical cyclone track accuracy and high spatial agreement for temperature and moisture extremes. However, a consistent subseasonal failure mode emerges: while large-scale circulation patterns remain moderately skillful at 14-21 day leads, threshold-based extreme intensity collapses as fields regress toward climatology. This divergence indicates that Aurora retains synoptic-scale dynamical structure but loses surface-impact amplitude beyond 7-10 days. The practical predictability horizon for deterministic AI extreme-event forecasting therefore remains constrained by intrinsic atmospheric dynamics.

Motivation & Objective

  • Assess the forecast skill of Aurora for selected weather extremes across multiple regimes.
  • Quantify performance across lead times from 1 to 21 days.
  • Identify failure modes in predictability, especially for surface-impact amplitudes at longer leads.

Proposed method

  • Use an event-based framework spanning tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation.
  • Evaluate forecast skill at lead times from 1 to 21 days.
  • Analyze both synoptic-scale structure and surface-impact amplitude to identify divergence from climatology.

Experimental results

Research questions

  • RQ1How does Aurora perform in predicting selected weather extremes at 1–7 days across different event types?
  • RQ2Does the model retain synoptic-scale dynamical structure at 14–21 day leads, and how does this relate to surface-impact amplitudes?
  • RQ3What are the failure modes of Aurora for extreme intensity forecasts beyond 7–10 days?

Key findings

  • Aurora shows strong short-range skill (1–7 days) across event types.
  • Tropical cyclone track accuracy is competitive and temperature/moisture extremes show high spatial agreement.
  • A consistent subseasonal failure mode emerges: threshold-based extreme intensity collapses as fields regress toward climatology at 14–21 days.
  • Aurora retains synoptic-scale dynamical structure but loses surface-impact amplitude beyond 7–10 days.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.