[Paper Review] Evaluation of Tropical Cyclone Track and Intensity Forecasts from Artificial Intelligence Weather Prediction (AIWP) Models
This study evaluates four open-source AI weather prediction models—FourCastNetv1, FourCastNetv2-small, GraphCast-operational, and Pangu-Weather—for tropical cyclone (TC) track and intensity forecasting using National Hurricane Center (NHC) verification standards. While AIWP models show track forecast accuracy comparable to top operational models, they exhibit substantial intensity forecast biases, underestimating storm intensity by up to 24 h, limiting their utility without bias correction.
In just the past few years multiple data-driven Artificial Intelligence Weather Prediction (AIWP) models have been developed, with new versions appearing almost monthly. Given this rapid development, the applicability of these models to operational forecasting has yet to be adequately explored and documented. To assess their utility for operational tropical cyclone (TC) forecasting, the NHC verification procedure is used to evaluate seven-day track and intensity predictions for northern hemisphere TCs from May-November 2023. Four open-source AIWP models are considered (FourCastNetv1, FourCastNetv2-small, GraphCast-operational and Pangu-Weather). The AIWP track forecast errors and detection rates are comparable to those from the best-performing operational forecast models. However, the AIWP intensity forecast errors are larger than those of even the simplest intensity forecasts based on climatology and persistence. The AIWP models almost always reduce the TC intensity, especially within the first 24 h of the forecast, resulting in a substantial low bias. The contribution of the AIWP models to the NHC model consensus was also evaluated. The consensus track errors are reduced by up to 11% at the longer time periods. The five-day NHC official track forecasts have improved by about 2% per year since 2001, so this represents more than a five-year gain in accuracy. Despite substantial negative intensity biases, the AIWP models have a neutral impact on the intensity consensus. These results show that the current formulation of the AIWP models have promise for operational TC track forecasts, but improved bias corrections or model reformulations will be needed for accurate intensity forecasts.
Motivation & Objective
- To assess the operational utility of emerging AI weather prediction (AIWP) models for tropical cyclone (TC) forecasting.
- To evaluate AIWP model performance in predicting TC track and intensity against NHC-verified standards.
- To determine the impact of AIWP models on the NHC official forecast consensus, particularly for track and intensity.
- To identify key limitations, especially in intensity forecasting, and suggest improvements for operational integration.
Proposed method
- The study uses the NHC verification procedure to evaluate seven-day track and intensity forecasts for Northern Hemisphere tropical cyclones from May to November 2023.
- Four open-source AIWP models—FourCastNetv1, FourCastNetv2-small, GraphCast-operational, and Pangu-Weather—are evaluated using historical atmospheric and oceanic reanalysis data as input.
- Track forecast errors and detection rates are computed and compared to operational models and climatology-persistence (C-P) baselines.
- Intensity forecast errors are quantified and analyzed for bias, particularly focusing on the first 24 hours of the forecast.
- The contribution of AIWP models to the NHC model consensus is assessed by measuring changes in consensus forecast error over time.
- Bias correction techniques are evaluated implicitly by analyzing forecast deviations from observed intensity trends.
Experimental results
Research questions
- RQ1How do AIWP models compare to operational models in predicting tropical cyclone track and intensity over a seven-day period?
- RQ2What is the magnitude and temporal evolution of intensity forecast bias in AIWP models, particularly within the first 24 hours?
- RQ3To what extent do AIWP models improve the NHC official forecast consensus, especially in track prediction?
- RQ4Why do AIWP models consistently underestimate tropical cyclone intensity despite strong track performance?
- RQ5Can the current AIWP model formulations be operationally integrated without significant bias correction?
Key findings
- AIWP models achieve track forecast errors comparable to the best operational models, with consensus track errors reduced by up to 11% at longer lead times.
- The five-day NHC official track forecast accuracy has improved by approximately 2% per year since 2001, and the AIWP contribution equates to more than five years of progress in track accuracy.
- AIWP models exhibit a substantial negative intensity bias, underestimating storm intensity—especially within the first 24 hours—resulting in larger errors than even simple climatology and persistence (C-P) forecasts.
- Despite the strong intensity bias, AIWP models have a neutral impact on the NHC intensity consensus, suggesting their errors do not degrade the consensus output.
- The study concludes that while AIWP models show strong promise for track forecasting, they require improved bias correction or model reformulation to be operationally viable for intensity prediction.
- The results highlight that current AIWP formulations are not yet suitable for direct use in intensity forecasting without significant post-processing or architectural refinement.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.