Skip to main content
QUICK REVIEW

[论文解读] Predicting Road Flooding Risk with Machine Learning Approaches Using Crowdsourced Reports and Fine-grained Traffic Data

Faxi Yuan, William H. Mobley|arXiv (Cornell University)|Aug 30, 2021
Flood Risk Assessment and Management参考文献 68被引用 5
一句话总结

本研究利用众包的 Waze 报告和细粒度交通数据,开发机器学习模型,以预测德克萨斯州 Harris 县在 2017 年哈维飓风和 2019 年热带风暴艾米莉达期间的路面洪水风险。随机森林模型在哈维飓风中的 AUC 得分为 0.860,在艾米莉达风暴中为 0.790,表现出高度的预测准确性和稳定性,其中降水和地形特征被识别为最重要的预测因子。

ABSTRACT

The objective of this study is to predict road flooding risks based on topographic, hydrologic, and temporal precipitation features using machine learning models. Predictive flood monitoring of road network flooding status plays an essential role in community hazard mitigation, preparedness, and response activities. Existing studies related to the estimation of road inundations either lack observed road inundation data for model validations or focus mainly on road inundation exposure assessment based on flood maps. This study addresses this limitation by using crowdsourced and fine-grained traffic data as an indicator of road inundation, and topographic, hydrologic, and temporal precipitation features as predictor variables. Two tree-based machine learning models (random forest and AdaBoost) were then tested and trained for predicting road inundations in the contexts of 2017 Hurricane Harvey and 2019 Tropical Storm Imelda in Harris County, Texas. The findings from Hurricane Harvey indicate that precipitation is the most important feature for predicting road inundation susceptibility, and that topographic features are more essential than hydrologic features for predicting road inundations in both storm cases. The random forest and AdaBoost models had relatively high AUC scores (0.860 and 0.810 for Harvey respectively and 0.790 and 0.720 for Imelda respectively) with the random forest model performing better in both cases. The random forest model showed stable performance for Harvey, while varying significantly for Imelda. This study advances the emerging field of smart flood resilience in terms of predictive flood risk mapping at the road level. For example, such models could help impacted communities and emergency management agencies develop better preparedness and response strategies with improved situational awareness of road inundation likelihood as an extreme weather event unfolds.

研究动机与目标

  • 通过利用实时众包报告和基于传感器的交通数据作为路面洪水的实际代理,弥补现有洪水风险模型中缺乏观测到的路面淹没数据的不足。
  • 通过将地形、水文和时间降水特征整合到机器学习模型中,提高路面淹没易感性的预测准确性。
  • 开发一种可扩展的、数据驱动的方法,实现实时道路级别的洪水风险监测,以支持应急管理与社区准备。
  • 证明使用异构大数据源——Waze 报告和 INRIX 交通数据——训练稳健的机器学习模型在城市洪水预测中的可行性。
  • 通过实现动态、可更新的风险制图,反映极端天气事件期间基础设施的实时脆弱性,推动智能洪水韧性建设。

提出的方法

  • 利用 Waze 众包报告(例如“路面洪水”警报)和 INRIX 细粒度交通速度数据作为实际路面淹没事件的指标。
  • 收集并整合地形特征(例如高程、坡度)、水文特征(例如不透水性、粗糙度)以及时间降水数据(例如风暴强度、持续时间)作为预测变量。
  • 在历史风暴事件(2017 年哈维飓风、2019 年热带风暴艾米莉达)上训练并评估两种基于树的机器学习模型——随机森林和 AdaBoost。
  • 通过特征重要性分析识别对路面洪水易感性最具影响力的预测因子。
  • 使用受试者工作特征曲线下面积(AUC)评估模型性能,并通过交叉验证确保稳健性。
  • 通过从交通数据集中随机抽样非洪水道路来解决数据不平衡问题,确保两次风暴事件的训练集均保持平衡。

实验结果

研究问题

  • RQ1众包的 Waze 报告和细粒度交通数据是否能有效作为机器学习模型中路面淹没事件的代理?
  • RQ2地形、水文和时间降水特征的何种组合在主要风暴事件中对路面洪水风险预测最具影响力?
  • RQ3树基机器学习模型(随机森林和 AdaBoost)在不同风暴事件中预测路面淹没的性能与稳定性如何比较?
  • RQ4降水和地形特征在预测路面洪水易感性方面,相较于水文特征的优越程度如何?
  • RQ5在数据需求较少的前提下,基于异质数据源(Waze 和 INRIX)训练的模型能否推广至其他地区和风暴情景?

主要发现

  • 随机森林模型在哈维飓风中的 AUC 得分为 0.860,在热带风暴艾米莉达中为 0.790,表明其具有出色的预测性能和稳健性。
  • 降水特征在两次风暴事件中均被识别为路面洪水风险最重要的预测因子,凸显了降雨强度和持续时间的主导作用。
  • 地形特征(例如高程、坡度)比水文特征(例如不透水性、粗糙度)更具影响力,表明地形在局部洪水中起着关键作用。
  • 随机森林模型在两次事件中均表现出稳定的性能,而 AdaBoost 模型则表现出更高的变异性,尤其在艾米莉达事件中(AUC 0.720)。
  • 尽管存在数据可用性限制以及土地覆盖数据时间错位(2016 年 NLDAS 数据用于 2017 年和 2019 年风暴),模型仍保持了较高的预测准确性,证明了利用现有数据实现实时风险建模的可行性。
  • 本研究证实,整合众包和基于传感器的数据可实现准确、可扩展且可更新的路面洪水风险预测,支持灾难响应中的实时态势感知。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。