[论文解读] Robust subsampling-based sparse Bayesian inference to tackle four challenges (large noise, outliers, data integration, and extrapolation) in the discovery of physical laws from data
本文提出了一种基于鲁棒子采样策略的稀疏贝叶斯推理方法,以在高噪声、异常值、跨实验数据整合以及外推四种挑战下从数据中发现物理规律。通过将最优子采样与稀疏贝叶斯学习相结合,该方法提升了准确性和泛化能力,在噪声大、含异常值及异质性数据集的数值基准测试中优于现有方法。
The derivation of physical laws is a dominant topic in scientific research. We propose a new method capable of discovering the physical laws from data to tackle four challenges in the previous methods. The four challenges are: (1) large noise in the data, (2) outliers in the data, (3) integrating the data collected from different experiments, and (4) extrapolating the solutions to the areas that have no available data. To resolve these four challenges, we try to discover the governing differential equations and develop a model-discovering method based on sparse Bayesian inference and subsampling. The subsampling technique is used for improving the accuracy of the Bayesian learning algorithm here, while it is usually employed for estimating statistics or speeding up algorithms elsewhere. The optimal subsampling size is moderate, neither too small nor too big. Another merit of our method is that it can work with limited data by the virtue of Bayesian inference. We demonstrate how to use our method to tackle the four aforementioned challenges step by step through numerical examples: (1) predator-prey model with noise, (2) shallow water equations with outliers, (3) heat diffusion with random initial and boundary conditions, and (4) fish-harvesting problem with bifurcations. Numerical results show that the robustness and accuracy of our new method is significantly better than the other model-discovering methods and traditional regression methods.
研究动机与目标
- 解决现有模型发现方法在处理实验数据中高噪声水平时的局限性。
- 克服异常值对从数据中发现物理规律准确性造成的不利影响。
- 实现从不同条件下的多种异质实验中收集的数据的有效整合。
- 将模型发现能力扩展至超出可用训练数据范围的区域。
- 开发一种数据高效的方法,在观测数据有限的情况下仍能保持高准确性,借助贝叶斯推理实现。
提出的方法
- 利用子采样提升贝叶斯学习的准确性,将子采样不仅视为加速技术,更作为鲁棒推理的核心组成部分。
- 采用稀疏贝叶斯推理识别控制微分方程中最相关的项,提升可解释性并减少过拟合。
- 优化子采样规模,使其适中——既不过小也不过大,以平衡统计可靠性与计算效率。
- 利用贝叶斯先验知识提升数据有限情况下的学习性能,实现在数据稀缺场景下的可靠推理。
- 通过迭代应用该方法,从噪声大或不完整的观测中估计系数与结构,以发现控制方程。
实验结果
研究问题
- RQ1如何使模型发现方法对观测数据中的大噪声具有鲁棒性?
- RQ2当训练数据中存在异常值时,该方法在多大程度上能保持准确性?
- RQ3是否存在一个统一框架,能够整合具有不同初始条件和边界条件的多个实验的数据?
- RQ4所发现的模型在超出观测数据范围的区域中,其外推性能如何?
- RQ5子采样在提升贝叶斯模型发现的鲁棒性与准确性方面,其最优作用是什么?
主要发现
- 所提出的方法在所有四项挑战下,其准确性和鲁棒性均显著优于传统回归方法和现有模型发现方法。
- 采用适中规模的子采样可提升贝叶斯学习的准确性,表明子采样并非单纯的计算技巧,而是实现鲁棒推理的关键机制。
- 在捕食者-猎物模型中,尽管噪声水平很高,该方法仍成功发现了正确的控制方程。
- 在浅水方程的案例中,即使数据被异常值污染,该方法仍保持高准确性。
- 在具有随机初始与边界条件的热扩散问题中,该方法在不同实验配置下均表现出良好的泛化能力。
- 在具有分岔现象的鱼类捕捞问题中,该方法被准确建模,展现出强大的外推能力,超越了训练数据范围。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。