Skip to main content
QUICK REVIEW

[论文解读] On sparse regression, Lp-regularization, and automated model discovery

Jeremy A. McCulloch, Skyler R. St. Pierre|arXiv (Cornell University)|Oct 9, 2023
Machine Learning in Materials Science被引用 6
一句话总结

本文提出一种混合神经网络方法,结合Lp正则化与物理约束,实现在本构建模中的自动、可解释的模型发现。结果表明,L0正则化通过实现透明、可微调的稀疏性,独特地平衡了可解释性与准确性,在保持模型保真度和减少偏差方面优于L1(Lasso)和L2(Ridge)。

ABSTRACT

Sparse regression and feature extraction are the cornerstones of knowledge discovery from massive data. Their goal is to discover interpretable and predictive models that provide simple relationships among scientific variables. While the statistical tools for model discovery are well established in the context of linear regression, their generalization to nonlinear regression in material modeling is highly problem-specific and insufficiently understood. Here we explore the potential of neural networks for automatic model discovery and induce sparsity by a hybrid approach that combines two strategies: regularization and physical constraints. We integrate the concept of Lp regularization for subset selection with constitutive neural networks that leverage our domain knowledge in kinematics and thermodynamics. We train our networks with both, synthetic and real data, and perform several thousand discovery runs to infer common guidelines and trends: L2 regularization or ridge regression is unsuitable for model discovery; L1 regularization or lasso promotes sparsity, but induces strong bias; only L0 regularization allows us to transparently fine-tune the trade-off between interpretability and predictability, simplicity and accuracy, and bias and variance. With these insights, we demonstrate that Lp regularized constitutive neural networks can simultaneously discover both, interpretable models and physically meaningful parameters. We anticipate that our findings will generalize to alternative discovery techniques such as sparse and symbolic regression, and to other domains such as biology, chemistry, or medicine. Our ability to automatically discover material models from data could have tremendous applications in generative material design and open new opportunities to manipulate matter, alter properties of existing materials, and discover new materials with user-defined properties.

研究动机与目标

  • 解决从固体力学数据中自动发现可解释、可预测的本构模型的挑战。
  • 研究不同Lp正则化策略(L0、L1、L2)在自动化发现过程中促进稀疏性与模型可解释性的有效性。
  • 将运动学与热力学的物理约束整合到神经网络中,以确保物理上合理的参数估计。
  • 开发一种透明、无偏差的发现框架,避免L1正则化引入的强偏差以及L2正则化缺乏稀疏性的问题。
  • 建立适用于材料建模之外的稀疏回归与模型发现的通用指导原则

提出的方法

  • 采用嵌入运动学与热力学领域知识的本构神经网络,以确保物理一致性。
  • 应用Lp正则化(L0、L1、L2)以诱导稀疏性,并在候选模型项中执行子集选择。
  • 采用自下而上的迭代增密策略,从最优单一项模型开始,逐步添加最相关的项,直至收敛。
  • 在数千次发现运行中,对合成数据和真实实验数据进行模型训练,以评估性能与稳定性。
  • 通过在用户定义的收敛标准下最小化损失函数,比较不同正则化策略,将模型复杂度限制在1–4项之间以保证可解释性。
  • 通过系统比较不同正则化类型下的模型精度、稀疏性、偏差与方差,验证结果
Figure 1: Lp regularization. Contours of regularization term, $L_{\rm{p}}=\alpha\,\sum_{i=1}^{n_{\rm{para}}}||\,\mbox{\boldmath$\theta$}{}\,||_{p}^{p}$ with $||\,\mbox{\boldmath$\theta$}{}\,||_{p}^{p}=|\,w_{i}\,|^{p}$ , for varying powers, $p=[0.25,0.5,0.75,1,1.5,2,4,8]$ , evaluated for two paramete
Figure 1: Lp regularization. Contours of regularization term, $L_{\rm{p}}=\alpha\,\sum_{i=1}^{n_{\rm{para}}}||\,\mbox{\boldmath$\theta$}{}\,||_{p}^{p}$ with $||\,\mbox{\boldmath$\theta$}{}\,||_{p}^{p}=|\,w_{i}\,|^{p}$ , for varying powers, $p=[0.25,0.5,0.75,1,1.5,2,4,8]$ , evaluated for two paramete

实验结果

研究问题

  • RQ1不同Lp正则化策略(L0、L1、L2)如何影响本构模型发现中可解释性与预测准确性的权衡?
  • RQ2L0正则化能否在不引入强偏差的情况下,实现对稀疏性的透明、可微调控制,从而避免L1正则化的缺陷?
  • RQ3将物理约束(如客观性、拟凸性)与神经网络结合,如何提升所发现模型的可靠性与可解释性?
  • RQ4在非凸、非线性回归问题中,自下而上的增密方法与自上而下的阈值法相比,对模型发现性能有何影响?
  • RQ5Lp-正则化神经网络在其他模型发现技术(如符号回归或稀疏回归)中,如在生物学、化学或医学领域,具有多大程度的泛化能力?

主要发现

  • L2正则化(岭回归)不适用于模型发现,因其无法诱导稀疏性,且保持所有参数,导致模型不可解释。
  • L1正则化(Lasso)虽促进稀疏性,但引入强偏差,常扭曲真实底层模型,降低预测准确性。
  • L0正则化通过允许无偏差的精确子集选择,唯一地实现了对可解释性与准确性的透明、可微调控制。
  • 所提出的自下而上的增密策略(基于损失改善逐步添加项)在具有多个局部极小值的非凸问题中,比自上而下的阈值法更有效。
  • 该方法成功从合成数据和真实数据中发现可解释、具有物理意义的本构模型,仅含1–4项的模型实现了高精度与低方差。
  • 该框架可泛化至本构建模之外,预计在符号回归、稀疏回归以及其它科学领域的自动化发现中同样有效
Figure 2: Invariant based neural network for automated model discovery. The network takes the deformation gradient $F$ as input and outputs the free energy function $\psi$ from which we calculate the stress $\mbox{\boldmath$P$}{}=\partial\psi/\partial\mbox{\boldmath$F$}{}$ . The network is invariant
Figure 2: Invariant based neural network for automated model discovery. The network takes the deformation gradient $F$ as input and outputs the free energy function $\psi$ from which we calculate the stress $\mbox{\boldmath$P$}{}=\partial\psi/\partial\mbox{\boldmath$F$}{}$ . The network is invariant

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。