Skip to main content
QUICK REVIEW

[论文解读] A Learning Framework for An Accurate Prediction of Rainfall Rates

Hamidreza Ghasemi Damavandi, Reepal Shah|arXiv (Cornell University)|Jan 17, 2019
Precipitation Measurement and Analysis参考文献 11被引用 3
一句话总结

本研究提出了一种基于梯度提升的机器学习框架,用于特征选择及多种回归模型(特别是随机森林)的月降雨率预测,以印度河盆地的气候变量为输入。随机森林模型在皮尔逊相关系数方面表现最佳,且平均绝对误差最低。相对湿度(300 mb,150 mb)、u-风(700 mb)和空气温度(150 mb,10 mb)被识别为最重要的预测特征。

ABSTRACT

The present work is aimed to examine the potential of advanced machine learning strategies to predict the monthly rainfall (precipitation) for the Indus Basin, using climatological variables such as air temperature, geo-potential height, relative humidity and elevation. In this work, the focus is on thirteen geographical locations, called index points, within the basin. Arguably, not all of the hydrological components are relevant to the precipitation rate, and therefore, need to be filtered out, leading to a lower-dimensional feature space. Towards this goal, we adopted the gradient boosting method to extract the most contributive features for precipitation rate prediction. Five state-of-the-art machine learning methods have then been trained where pearson correlation coefficient and mean absolute error have been reported as the prediction performance criteria. The Random Forest regression model outperformed the other regression models achieving the maximum pearson correlation coefficient and minimum mean absolute error for most of the index points. Our results suggest the relative humidity (for pressure levels of 300 mb and 150 mb, respectively), the u-direction wind (for pressure level of 700 mb), air temperature (for pressure levels of 150 mb and 10 mb, respectively) as the top five influencing features for accurate forecasting the precipitation rate.

研究动机与目标

  • 开发一种数据驱动的机器学习框架,以实现印度河盆地月降雨率的高精度预测。
  • 从温度、湿度、风速和位势高度等气候变量中识别出最具影响力的水文预测因子。
  • 使用皮尔逊相关系数和平均绝对误差作为评估指标,比较五种先进机器学习模型的性能。
  • 通过基于梯度提升的特征选择方法,过滤无关的水文成分,降低特征维度。
  • 为区域降雨预测提供一种计算效率更高的替代方案,以取代传统的物理驱动水文模型。

提出的方法

  • 使用CHIRPS v2降水数据(0.05°分辨率)和NCEP-NCAR再分析数据(1981–2017年,2.5°分辨率)作为大气预测因子。
  • 通过将不同气压层上的水文变量(如相对湿度、风速分量)视为独立预测因子,共提取出85个独立特征。
  • 应用梯度提升方法,在印度河盆地13个指标点上对降雨预测最具贡献的特征进行排序与提取。
  • 在90%的数据上训练五种机器学习模型(随机森林、XGBoost、LightGBM、SVM和ANN),10%用于测试。
  • 通过交叉验证优化超参数,并使用皮尔逊相关系数和平均绝对误差(MAE)评估模型性能。
  • 开展特征重要性分析,基于各指标点上特征被选择的频率,识别出最重要的预测因子。

实验结果

研究问题

  • RQ1在使用气候变量预测印度河盆地月降雨率时,哪种机器学习模型表现最佳?
  • RQ2在不同大气气压层上,温度、湿度、风速和位势高度等水文特征中,哪些对降水具有最强预测力?
  • RQ3梯度提升在减少特征空间并提升降雨率预测准确性方面有多高效?
  • RQ4特定大气变量(如300 mb处的相对湿度、700 mb处的u-风)对准确降雨预测的相对贡献如何?
  • RQ5在区域降雨预测中,数据驱动模型能否在准确性和计算效率方面超越传统物理驱动模型?

主要发现

  • 随机森林回归模型在13个指标点中的大多数区域表现最优,皮尔逊相关系数最高,平均绝对误差最低。
  • 300 mb和150 mb处的相对湿度被频繁选为最优先预测因子,分别在11个和8个模型中位列第一。
  • 700 mb处的u-风在8个指标点中被选为关键特征,表明其具有显著的预测相关性。
  • 150 mb和10 mb处的空气温度位列前五位最具影响力特征,分别在7个和6个模型中被选中。
  • 特征选择频率分析显示,rh^l1(1000 mb处的相对湿度)是最具主导性的预测因子,在11个指标点中出现。
  • 在多个气压层上,相对湿度、u-风和空气温度的组合持续构成最具有信息量的特征集合,对准确降雨预测具有关键作用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。