[Paper Review] Time Series Predictions in Unmonitored Sites: A Survey of Machine Learning Techniques in Water Resources
This survey reviews machine learning techniques for predicting hydrological and water quality time series in unmonitored basins, emphasizing deep learning, transfer learning, and knowledge-guided models. It identifies key challenges in data scarcity, model interpretability, and transferability, and highlights the need for explainable AI and integration of hydrological process knowledge to improve prediction accuracy and reliability in data-scarce regions.
Prediction of dynamic environmental variables in unmonitored sites remains a long-standing challenge for water resources science. The majority of the world's freshwater resources have inadequate monitoring of critical environmental variables needed for management. Yet, the need to have widespread predictions of hydrological variables such as river flow and water quality has become increasingly urgent due to climate and land use change over the past decades, and their associated impacts on water resources. Modern machine learning methods increasingly outperform their process-based and empirical model counterparts for hydrologic time series prediction with their ability to extract information from large, diverse data sets. We review relevant state-of-the art applications of machine learning for streamflow, water quality, and other water resources prediction and discuss opportunities to improve the use of machine learning with emerging methods for incorporating watershed characteristics into deep learning models, transfer learning, and incorporating process knowledge into machine learning models. The analysis here suggests most prior efforts have been focused on deep learning learning frameworks built on many sites for predictions at daily time scales in the United States, but that comparisons between different classes of machine learning methods are few and inadequate. We identify several open questions for time series predictions in unmonitored sites that include incorporating dynamic inputs and site characteristics, mechanistic understanding and spatial context, and explainable AI techniques in modern machine learning frameworks.
Motivation & Objective
- To address the critical gap in predicting hydrological variables in unmonitored basins due to sparse observational data.
- To synthesize state-of-the-art machine learning methods for streamflow, water quality, and environmental variable prediction.
- To evaluate the strengths and limitations of different ML approaches, especially in comparison to process-based models.
- To identify open research questions in data efficiency, model explainability, and integration of hydrological process knowledge.
- To guide future research by outlining best practices and interdisciplinary collaboration needs in water resources ML.
Proposed method
- Systematic review of 230+ studies on machine learning applications in ungauged basins (PUBs) from 2010 to 2022.
- Categorization of ML techniques into deep learning (LSTM, GNNs, TCNs), transfer learning, and knowledge-guided models (e.g., differentiable process-based models).
- Analysis of model performance across time scales (daily predictions dominant), geographic focus (primarily U.S.), and data types (streamflow, water quality, lake temperature).
- Evaluation of interpretability methods such as SHAP, integrated gradients, and layerwise relevance propagation for explaining model decisions.
- Integration of watershed characteristics and dynamic inputs (e.g., meteorological data) into deep learning frameworks to improve generalization.
- Comparison of classical ML (e.g., random forest, XGBoost) with deep learning models in terms of accuracy, interpretability, and data requirements.

Experimental results
Research questions
- RQ1How do different machine learning models (e.g., LSTM, XGBoost, GNNs) compare in predicting streamflow and water quality in unmonitored basins?
- RQ2To what extent can transfer learning improve model generalization across ungauged catchments with limited training data?
- RQ3How can process-based hydrological knowledge be embedded into deep learning models to improve interpretability and physical consistency?
- RQ4What are the key challenges in model explainability (XAI) for ML in PUBs, and how can methods like SHAP and integrated gradients enhance trust and usability?
- RQ5How can dynamic inputs (e.g., rainfall, temperature) and spatial context be effectively incorporated into ML models for better spatiotemporal predictions?
Key findings
- Most studies focus on daily-scale predictions in the United States, with limited geographic and temporal diversity in training data.
- Deep learning models, particularly LSTMs and GNNs, outperform traditional empirical and process-based models in prediction accuracy when sufficient data are available.
- Transfer learning shows promise in improving performance in unmonitored basins by leveraging pre-trained models from monitored regions.
- Classical ML models like XGBoost and random forest remain widely used due to their interpretability and robustness with small datasets.
- Explainable AI (XAI) techniques such as SHAP and integrated gradients can reveal temporal and spatial attention patterns in models like LSTMs and GNNs, enhancing model transparency.
- Despite progress, systematic comparisons between ML model classes are rare, and there is a lack of standardized benchmarks for PUBs, especially for water quality and long-term predictions.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.