[论文解读] A mixed model approach to drought prediction using artificial neural networks: Case of an operational drought monitoring environment
本研究提出一种混合模型方法,结合广义可加模型(GAM)与人工神经网络(ANN),利用滞后降水和植被指数数据,提升未来1个月的干旱预测性能。通过将102个GAM模型筛选至R² > 0.7的21个模型,该方法降低了模型空间的复杂度,并识别出最优预测变量,最终得到的ANN冠军模型在样本外数据上实现R² = 0.78,表现出优异的预测性能,适用于业务化干旱监测。
Droughts, with their increasing frequency of occurrence, continue to negatively affect livelihoods and elements at risk. For example, the 2011 in drought in east Africa has caused massive losses document to have cost the Kenyan economy over $12bn. With the foregoing, the demand for ex-ante drought monitoring systems is ever-increasing. The study uses 10 precipitation and vegetation variables that are lagged over 1, 2 and 3-month time-steps to predict drought situations. In the model space search for the most predictive artificial neural network (ANN) model, as opposed to the traditional greedy search for the most predictive variables, we use the General Additive Model (GAM) approach. Together with a set of assumptions, we thereby reduce the cardinality of the space of models. Even though we build a total of 102 GAM models, only 21 have R2 greater than 0.7 and are thus subjected to the ANN process. The ANN process itself uses the brute-force approach that automatically partitions the training data into 10 sub-samples, builds the ANN models in these samples and evaluates their performance using multiple metrics. The results show the superiority of 1-month lag of the variables as compared to longer time lags of 2 and 3 months. The champion ANN model recorded an R2 of 0.78 in model testing using the out-of-sample data. This illustrates its ability to be a good predictor of drought situations 1-month ahead. Investigated as a classifier, the champion has a modest accuracy of 66% and a multi-class area under the ROC curve (AUROC) of 89.99%
研究动机与目标
- 为应对日益增长的对高精度、早期干旱预测的需求,以减轻社会经济影响。
- 通过使用GAM预先筛选信息丰富的预测变量,降低干旱预测中高维模型空间的复杂度。
- 基于滞后气象与遥感变量,开发一种稳健、数据驱动的ANN模型,用于业务化干旱监测。
- 通过样本外测试和多分类评估指标,评估最终ANN模型的预测性能。
提出的方法
- 本研究使用10个滞后变量(降水和植被指数的1、2、3个月滞后)作为输入特征。
- 应用广义可加模型(GAM)探索预测变量与干旱之间的关系,将模型空间缩减至21个高性能模型(R² > 0.7)。
- 利用筛选出的21个GAM模型,通过10折划分的暴力交叉验证方法,训练并评估人工神经网络(ANN)。
- ANN训练过程采用多种性能指标,评估模型在未见数据上的泛化能力与鲁棒性。
- 模型选择优先考虑在独立测试集上具有最高样本外R²的ANN。
- 最终模型还通过准确率和受试者工作特征曲线下面积(AUROC)作为分类器进行评估。
实验结果
研究问题
- RQ1与传统变量选择方法相比,混合GAM-ANN方法是否能提升干旱预测的准确性?
- RQ2气象与植被变量的哪个滞后时间步长(1、2或3个月)能产生最强的预测性能?
- RQ3所选ANN模型在样本外数据上的预测性能如何,用于1个月后干旱的预测?
- RQ4冠军模型作为多分类干旱分类器的性能如何?
主要发现
- 降水和植被变量的1个月滞后时间步长表现出最强的预测性能,优于2和3个月滞后。
- 冠军ANN模型在样本外数据上实现了R² = 0.78,表明其在1个月后干旱预测方面具有强大的预测能力。
- 在测试的102个GAM模型中,仅有21个达到R² > 0.7,证明GAM在过滤模型空间方面的有效性。
- 作为分类器,冠军模型实现了66%的适度准确率和89.99%的多分类AUROC,表明其在不同干旱严重程度类别间具有良好的判别能力。
- 本研究证实,混合模型方法——即使用GAM进行变量筛选,再利用ANN进行最终建模——显著提升了业务化干旱监测中的预测性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。