Skip to main content
QUICK REVIEW

[论文解读] Time Series Predictions in Unmonitored Sites: A Survey of Machine Learning Techniques in Water Resources

Jared Willard, Charuleka Varadharajan|arXiv (Cornell University)|Aug 18, 2023
Hydrological Forecasting Using AI被引用 4
一句话总结

本综述回顾了用于预测未监测流域水文和水质时间序列的机器学习技术,重点聚焦于深度学习、迁移学习和知识引导模型。它识别出数据稀缺、模型可解释性和可迁移性等关键挑战,并强调需要采用可解释人工智能(XAI)并整合水文过程知识,以提升数据匮乏地区预测的准确性和可靠性。

ABSTRACT

Prediction of dynamic environmental variables in unmonitored sites remains a long-standing challenge for water resources science. The majority of the world's freshwater resources have inadequate monitoring of critical environmental variables needed for management. Yet, the need to have widespread predictions of hydrological variables such as river flow and water quality has become increasingly urgent due to climate and land use change over the past decades, and their associated impacts on water resources. Modern machine learning methods increasingly outperform their process-based and empirical model counterparts for hydrologic time series prediction with their ability to extract information from large, diverse data sets. We review relevant state-of-the art applications of machine learning for streamflow, water quality, and other water resources prediction and discuss opportunities to improve the use of machine learning with emerging methods for incorporating watershed characteristics into deep learning models, transfer learning, and incorporating process knowledge into machine learning models. The analysis here suggests most prior efforts have been focused on deep learning learning frameworks built on many sites for predictions at daily time scales in the United States, but that comparisons between different classes of machine learning methods are few and inadequate. We identify several open questions for time series predictions in unmonitored sites that include incorporating dynamic inputs and site characteristics, mechanistic understanding and spatial context, and explainable AI techniques in modern machine learning frameworks.

研究动机与目标

  • 为解决由于观测数据稀疏而导致的未监测流域水文变量预测关键缺口。
  • 综合梳理用于径流、水质及环境变量预测的最先进机器学习方法。
  • 评估不同机器学习方法的优势与局限性,特别是与过程模型的对比。
  • 识别在数据效率、模型可解释性以及水文过程知识整合方面的开放性研究问题。
  • 通过提出最佳实践和跨学科合作需求,为未来水文资源机器学习研究提供指导。

提出的方法

  • 系统性回顾2010至2022年间关于机器学习在无流量监测流域(PUBs)应用的230余项研究。
  • 将机器学习技术分类为深度学习(LSTM、图神经网络GNNs、时间卷积网络TCNs)、迁移学习以及知识引导模型(如可微分过程模型)。
  • 分析模型在时间尺度(以日尺度预测为主)、地理重点(主要集中在美国)和数据类型(径流、水质、湖泊温度)上的表现。
  • 评估SHAP、积分梯度和逐层重要性传播等可解释性方法,用于解释模型决策过程。
  • 将流域特征和动态输入(如气象数据)整合进深度学习框架,以提升模型泛化能力。
  • 从准确性、可解释性和数据需求角度,对比经典机器学习(如随机森林、XGBoost)与深度学习模型的性能。
Figure 1: Example of an long short-term memory (LSTM) network model with directly concatenated site characteristics and dynamic inputs
Figure 1: Example of an long short-term memory (LSTM) network model with directly concatenated site characteristics and dynamic inputs

实验结果

研究问题

  • RQ1在未监测流域中,不同机器学习模型(如LSTM、XGBoost、GNNs)在预测径流和水质方面表现如何比较?
  • RQ2在训练数据有限的无流量监测流域中,迁移学习在多大程度上能提升模型泛化能力?
  • RQ3如何将基于过程的水文知识嵌入深度学习模型中,以提升可解释性和物理一致性?
  • RQ4在PUBs中,机器学习的可解释性(XAI)面临哪些关键挑战?SHAP和积分梯度等方法如何增强信任度和可用性?
  • RQ5如何有效整合动态输入(如降雨、气温)和空间上下文信息,以提升机器学习模型的时空预测能力?

主要发现

  • 大多数研究聚焦于美国的每日尺度预测,训练数据在地理和时间维度上的多样性有限。
  • 当数据充足时,深度学习模型(尤其是LSTM和GNNs)在预测精度上优于传统经验模型和过程模型。
  • 迁移学习在利用已监测区域的预训练模型方面展现出提升无流量监测流域性能的潜力。
  • 由于其可解释性和在小样本数据上的稳健性,XGBoost和随机森林等经典机器学习模型仍被广泛使用。
  • 可解释人工智能(XAI)技术(如SHAP和积分梯度)可揭示LSTM和GNNs等模型中的时间与空间注意力模式,增强模型透明度。
  • 尽管已有进展,但不同机器学习模型类别之间的系统性比较仍较少,且缺乏针对PUBs的标准化基准,尤其在水质和长期预测方面。
Figure 2: Example of a combination static feature encoder neural network with a long short-term memory (LSTM) network model
Figure 2: Example of a combination static feature encoder neural network with a long short-term memory (LSTM) network model

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。