[Paper Review] Efficient surrogate modeling methods for large-scale Earth system models based on machine learning techniques
This paper proposes a machine learning-based surrogate modeling framework that uses singular value decomposition (SVD) for dimensionality reduction and Bayesian optimization to train a neural network surrogate from only 20 expensive Earth system model (ESM) simulations. The method achieves high accuracy—0.93 correlation and 0.02 mean squared error—across 42,660 carbon flux outputs, enabling fast, reusable predictions without retraining for new parameters or timeframes.
Improving predictive understanding of Earth system variability and change requires data-model integration. Efficient data-model integration for complex models requires surrogate modeling to reduce model evaluation time. However, building a surrogate of a large-scale Earth system model (ESM) with many output variables is computationally intensive because it involves a large number of expensive ESM simulations. In this effort, we propose an efficient surrogate method capable of using a few ESM runs to build an accurate and fast-to-evaluate surrogate system of model outputs over large spatial and temporal domains. We first use singular value decomposition to reduce the output dimensions, and then use Bayesian optimization techniques to generate an accurate neural network surrogate model based on limited ESM simulation samples. Our machine learning based surrogate methods can build and evaluate a large surrogate system of many variables quickly. Thus, whenever the quantities of interest change such as a different objective function, a new site, and a longer simulation time, we can simply extract the information of interest from the surrogate system without rebuilding new surrogates, which significantly saves computational efforts. We apply the proposed method to a regional ecosystem model to approximate the relationship between 8 model parameters and 42660 carbon flux outputs. Results indicate that using only 20 model simulations, we can build an accurate surrogate system of the 42660 variables, where the consistency between the surrogate prediction and actual model simulation is 0.93 and the mean squared error is 0.02. This highly-accurate and fast-to-evaluate surrogate system will greatly enhance the computational efficiency in data-model integration to improve predictions and advance our understanding of the Earth system.
Motivation & Objective
- To reduce the computational burden of data-model integration in large-scale Earth system models (ESMs) by creating fast, accurate surrogate models.
- To address the challenge of high-dimensional ESM outputs with tens of thousands of variables across spatial and temporal domains.
- To develop a reusable surrogate system that supports rapid re-evaluation for new parameters, objectives, or simulation durations without retraining.
- To enable efficient exploration of model parameter spaces for improved predictive understanding of Earth system variability and change.
Proposed method
- Apply singular value decomposition (SVD) to reduce the dimensionality of high-dimensional ESM output data, capturing dominant patterns in the output space.
- Use Bayesian optimization to intelligently select the minimal set of ESM simulations needed for training the surrogate model.
- Train a neural network surrogate on the reduced-dimensional output space using the selected simulation samples to ensure high accuracy with minimal data.
- Construct a single, unified surrogate system that can be queried for any subset of outputs or parameters without rebuilding.
- Leverage the low-rank structure of the output data to maintain computational efficiency and scalability across large spatial and temporal domains.
- Enable rapid re-evaluation of the surrogate for new objectives or parameters by directly extracting relevant outputs from the pre-trained system.
Experimental results
Research questions
- RQ1Can a surrogate model be constructed from only a small number of expensive ESM simulations while maintaining high accuracy across thousands of output variables?
- RQ2How effective is the combination of SVD and Bayesian optimization in reducing computational cost while preserving predictive fidelity in large-scale ESMs?
- RQ3To what extent can a single surrogate system support multiple queries—such as new parameters, objectives, or timeframes—without retraining?
- RQ4What level of accuracy can be achieved in predicting high-dimensional carbon flux outputs using limited ESM simulation data?
Key findings
- The surrogate model achieved a correlation of 0.93 between predicted and actual ESM outputs across 42,660 carbon flux variables using only 20 ESM simulations.
- The mean squared error (MSE) of the surrogate predictions was 0.02, indicating high predictive accuracy despite limited training data.
- The surrogate system enabled fast evaluation and reuse for new parameters or objectives without retraining, significantly reducing computational overhead.
- The SVD-based dimensionality reduction effectively captured the dominant modes of variability in the high-dimensional output space.
- Bayesian optimization enabled efficient sampling of the parameter space, minimizing the number of required ESM runs for high-fidelity surrogate construction.
- The method demonstrated scalability and robustness for large-scale Earth system modeling applications with complex, multi-variable outputs.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.