[Paper Review] Joint inversion of Time-Lapse Surface Gravity and Seismic Data for Monitoring of 3D CO$_2$ Plumes via Deep Learning
This paper proposes a novel 3D deep learning-based joint inversion framework that simultaneously processes time-lapse surface gravity and seismic data to reconstruct high-resolution subsurface CO₂ plume distributions. By leveraging a physics-simulated dataset from the Kimberlina site, the method outperforms gravity-only and seismic-only models in density and velocity reconstruction, segmentation accuracy, and R-squared metrics, demonstrating the effectiveness of multi-physics data fusion for carbon storage monitoring.
We introduce a fully 3D, deep learning-based approach for the joint inversion of time-lapse surface gravity and seismic data for reconstructing subsurface density and velocity models. The target application of this proposed inversion approach is the prediction of subsurface CO2 plumes as a complementary tool for monitoring CO2 sequestration deployments. Our joint inversion technique outperforms deep learning-based gravity-only and seismic-only inversion models, achieving improved density and velocity reconstruction, accurate segmentation, and higher R-squared coefficients. These results indicate that deep learning-based joint inversion is an effective tool for CO$_2$ storage monitoring. Future work will focus on validating our approach with larger datasets, simulations with other geological storage sites, and ultimately field data.
Motivation & Objective
- To develop a scalable, 3D deep learning-based inversion method for monitoring subsurface CO₂ plumes using time-lapse surface gravity and seismic data.
- To overcome the limitations of single-modality inversion by fusing complementary geophysical data for improved subsurface model resolution and accuracy.
- To validate the proposed joint inversion framework on realistic, physics-simulated CO₂ storage scenarios, particularly for geological carbon sequestration applications.
- To reduce computational cost and model complexity through architectural innovations like PocketNet and trilinear upsampling while maintaining high performance.
- To establish a foundation for future field validation and generalization to other storage sites such as Snohvit and Kimberlina.
Proposed method
- A 3D U-Net-based deep learning architecture is trained end-to-end on synthetic time-lapse gravity and seismic data derived from physics simulations of CO₂ injection at the Kimberlina site.
- The model employs the PocketNet architecture to reduce parameter count from ~33M to ~349K, significantly lowering training time and memory usage.
- Transposed convolutions are replaced with trilinear upsampling to improve inference efficiency and stability.
- A joint loss function combines data misfit and regularization terms, with future work exploring a Dice score-based coupling term to enforce spatial consistency between plume predictions.
- The method is trained in a supervised manner using ground-truth density and velocity models generated from reservoir simulations.
- Training is performed on 4 A100 GPUs with a batch size of 8, achieving convergence in ~400 epochs.

Experimental results
Research questions
- RQ1Can a deep learning-based joint inversion of surface gravity and seismic data outperform single-modality inversion in reconstructing 3D CO₂ plume distributions?
- RQ2How does the fusion of gravity and seismic data improve the accuracy of subsurface density and velocity models compared to individual data types?
- RQ3To what extent does architectural optimization, such as PocketNet and trilinear upsampling, reduce training cost without sacrificing model performance?
- RQ4How do the joint inversion results compare across key metrics such as R-squared, segmentation accuracy, and model misfit?
- RQ5Can the proposed method generalize to other geological storage sites beyond the Kimberlina simulation?
Key findings
- The joint inversion model achieves superior performance in density and velocity reconstruction compared to gravity-only and seismic-only models, with improved R-squared coefficients and lower misfit.
- Visual comparisons show that the joint model produces more accurate and spatially consistent plume segmentation than single-modality models.
- The joint model converges in approximately 400 epochs, compared to 200 for single-modality models, due to the increased complexity of learning multi-physics relationships.
- Training time per epoch is ~45 seconds on 4 A100 GPUs, which is higher than gravity-only (~20s) and seismic-only (~30s) models, but still computationally feasible.
- Inference inference time is under one second per prediction on a single A100 GPU, enabling real-time application potential.
- The use of PocketNet reduces model parameters by ~98.7%, significantly lowering memory and compute demands while maintaining high performance.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.