Skip to main content
QUICK REVIEW

[论文解读] A new framework for experimental design using Bayesian Evidential Learning: the case of wellhead protection area

Robin Thibaut, Eric Laloy|arXiv (Cornell University)|May 12, 2021
Groundwater flow and contamination studies参考文献 67被引用 33
一句话总结

该论文提出了一种贝叶斯证据学习(BEL)框架,可直接将示踪剂突破曲线与水源地保护区(WHPA)预测关联起来,实现快速、无需校准的随机不确定性量化及最优实验设计。通过在400个正向模拟实现上进行训练,BEL能够预测完整的后验WHPA分布,并识别出最具信息量的示踪剂注入井位置——该方法通过250个样本的测试集进行k折交叉验证得到验证,显著降低了计算成本,同时保持了高精度。

ABSTRACT

In this contribution, we predict the wellhead protection area (WHPA, target), the shape and extent of which is influenced by the distribution of hydraulic conductivity (K), from a small number of tracing experiments (predictor). Our first objective is to make stochastic predictions of the WHPA within the Bayesian Evidential Learning (BEL) framework, which aims to find a direct relationship between predictor and target using machine learning. This relationship is learned from a small set of training models (400) sampled from the prior distribution of K. The associated 400 pairs of simulated predictors and targets are obtained through forward modelling. Newly collected field data can then be directly used to predict the approximate posterior distribution of the corresponding WHPA. The uncertainty range of the posterior WHPA distribution is affected by the number and position of data sources (injection wells). Our second objective is to extend BEL to identify the optimal design of data source locations that minimizes the posterior uncertainty of the WHPA. This can be done explicitly, without averaging or approximating because once trained, the BEL model allows the computation of the posterior uncertainty corresponding to any new input data. We use the Modified Hausdorff Distance and the Structural Similarity index metrics to estimate the posterior uncertainty range of the WHPA. Increasing the number of injection wells effectively reduces the derived posterior WHPA uncertainty. Our approach can also estimate which injection wells are more informative than others, as validated through a k-fold cross-validation procedure. Overall, the application of BEL to experimental design makes it possible to identify the data sources maximizing the information content of any measurement data.

研究动机与目标

  • 开发一种计算高效的随机方法,用于在地下不确定性条件下预测水源地保护区(WHPA)。
  • 应用贝叶斯证据学习(BEL)以绕过传统模型校准,直接将示踪剂数据(预测变量)与WHPA(目标变量)预测关联。
  • 识别能最小化后验WHPA不确定性的最优示踪剂注入井位置。
  • 通过k折交叉验证与不确定性度量方法验证特定数据源的信息量。
  • 证明小规模训练集(400个模型)足以实现可靠的WHPA预测与实验设计。

提出的方法

  • 在400个水力传导率场的正向模拟实现及其对应的示踪剂突破曲线与WHPA形态上训练BEL模型。
  • 在降维空间中使用典型相关分析(CCA)学习突破曲线(预测变量)与WHPA(目标变量)之间的直接非线性映射。
  • 将训练好的BEL模型应用于新野外数据,直接计算WHPA的完整后验分布,无需迭代反演。
  • 使用修正的豪斯多夫距离(MHD)和结构相似性(SSIM)指数作为数据效用函数,量化后验不确定性。
  • 通过不同测试集大小(100和250个样本)的k折交叉验证,验证模型鲁棒性并确定最优数据集大小。
  • 通过比较各折与不同数据配置下的MHD与SSIM度量,评估单个注入井的信息含量。

实验结果

研究问题

  • RQ1贝叶斯证据学习(BEL)能否在无需模型校准或反演的情况下,仅从示踪剂突破曲线预测水源地保护区(WHPA)的完整后验分布?
  • RQ2哪些注入井位置能提供最高信息量,以减少WHPA预测不确定性?
  • RQ3为在WHPA实验设计中可靠地对数据源信息量进行排序,所需的最小测试样本数是多少?
  • RQ4训练数据集规模(如400个模型与1000个模型)如何影响WHPA预测与实验设计结果的鲁棒性?
  • RQ5与贝叶斯模型平均或代理建模等传统方法相比,基于BEL的实验设计在计算效率与准确性方面是否更具优势?

主要发现

  • 400个模型的训练集足以实现准确的WHPA预测与稳健的实验设计。
  • 最具信息量的注入井始终被排在第4、5和6号(下游井),其中第6号井在所有k折分割中均表现出最窄的不确定性区间。
  • 第1号井(上游井)始终信息量最低,不确定性区间最宽,表明其信息含量较低。
  • 为在k折交叉验证中实现一致且可靠的对数据源信息量的排序,至少需要250个测试样本。
  • 使用MHD与SSIM作为数据效用函数,可实现无需完整后验采样的直接、计算高效的不确定性量化。
  • BEL框架避免了马尔可夫链蒙特卡洛或代理建模的高计算成本,同时保持了预测精度,并支持显式实验设计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。