[Paper Review] Evaluating the Predictability of Selected Weather Extremes with Aurora, an AI Weather Forecast Model
The paper evaluates Aurora, an AI weather forecast model, for predictability of selected weather extremes across tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation at lead times from 1 to 21 days, highlighting strong short-range skill but a subseasonal failure mode in extreme intensity.
AI weather foundation models now achieve forecast skill comparable to numerical weather prediction at far lower computational cost, yet their predictability for high-impact extremes across dynamical regimes remains uncertain. We evaluate Aurora using an event-based framework spanning tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation at lead times from 1 to 21 days. Aurora demonstrates strong short-range (1-7 day) skill across event types, including competitive tropical cyclone track accuracy and high spatial agreement for temperature and moisture extremes. However, a consistent subseasonal failure mode emerges: while large-scale circulation patterns remain moderately skillful at 14-21 day leads, threshold-based extreme intensity collapses as fields regress toward climatology. This divergence indicates that Aurora retains synoptic-scale dynamical structure but loses surface-impact amplitude beyond 7-10 days. The practical predictability horizon for deterministic AI extreme-event forecasting therefore remains constrained by intrinsic atmospheric dynamics.
Motivation & Objective
- Assess the forecast skill of Aurora for selected weather extremes across multiple regimes.
- Quantify performance across lead times from 1 to 21 days.
- Identify failure modes in predictability, especially for surface-impact amplitudes at longer leads.
Proposed method
- Use an event-based framework spanning tropical cyclones, freezes, heatwaves, atmospheric rivers, and extreme precipitation.
- Evaluate forecast skill at lead times from 1 to 21 days.
- Analyze both synoptic-scale structure and surface-impact amplitude to identify divergence from climatology.
Experimental results
Research questions
- RQ1How does Aurora perform in predicting selected weather extremes at 1–7 days across different event types?
- RQ2Does the model retain synoptic-scale dynamical structure at 14–21 day leads, and how does this relate to surface-impact amplitudes?
- RQ3What are the failure modes of Aurora for extreme intensity forecasts beyond 7–10 days?
Key findings
- Aurora shows strong short-range skill (1–7 days) across event types.
- Tropical cyclone track accuracy is competitive and temperature/moisture extremes show high spatial agreement.
- A consistent subseasonal failure mode emerges: threshold-based extreme intensity collapses as fields regress toward climatology at 14–21 days.
- Aurora retains synoptic-scale dynamical structure but loses surface-impact amplitude beyond 7–10 days.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.