[Paper Review] Extreme sparsification of physics-augmented neural networks for interpretable model discovery in mechanics
This paper proposes extreme sparsification of physics-augmented neural networks using smoothed L⁰-regularization to discover interpretable, trustworthy constitutive models in solid mechanics. By jointly training for data fit and parameter sparsity while enforcing thermodynamic consistency, the method yields compact, human-readable functional forms for hyperelasticity, yield functions, and hardening laws—demonstrated successfully on synthetic and experimental data with high accuracy and strong extrapolation.
Data-driven constitutive modeling with neural networks has received increased interest in recent years due to its ability to easily incorporate physical and mechanistic constraints and to overcome the challenging and time-consuming task of formulating phenomenological constitutive laws that can accurately capture the observed material response. However, even though neural network-based constitutive laws have been shown to generalize proficiently, the generated representations are not easily interpretable due to their high number of trainable parameters. Sparse regression approaches exist that allow to obtaining interpretable expressions, but the user is tasked with creating a library of model forms which by construction limits their expressiveness to the functional forms provided in the libraries. In this work, we propose to train regularized physics-augmented neural network-based constitutive models utilizing a smoothed version of $L^{0}$-regularization. This aims to maintain the trustworthiness inherited by the physical constraints, but also enables interpretability which has not been possible thus far on any type of machine learning-based constitutive model where model forms were not assumed a-priory but were actually discovered. During the training process, the network simultaneously fits the training data and penalizes the number of active parameters, while also ensuring constitutive constraints such as thermodynamic consistency. We show that the method can reliably obtain interpretable and trustworthy constitutive models for compressible and incompressible hyperelasticity, yield functions, and hardening models for elastoplasticity, for synthetic and experimental data.
Motivation & Objective
- To overcome the lack of interpretability in data-driven neural network constitutive models despite their high expressiveness and generalization capability.
- To eliminate the need for pre-defined model form libraries in sparse regression approaches that limit functional expressiveness.
- To enable trustworthy, extrapolative constitutive modeling by combining physical constraints with parameter sparsity through regularization.
- To automate the discovery of interpretable, physically consistent constitutive laws from limited experimental or synthetic data.
- To bridge the gap between high-capacity neural networks and interpretable, human-readable mathematical expressions in solid mechanics modeling.
Proposed method
- The method employs a smoothed approximation of L⁰-regularization to penalize the number of active parameters during training, promoting extreme sparsity.
- Physics-augmented neural networks are trained to simultaneously fit training data and minimize the number of non-zero parameters, ensuring model simplicity.
- Thermodynamic consistency, objectivity, and material symmetry are enforced during network architecture and loss function design.
- The approach uses a single nonlinear activation function per neuron, with pruning guided by the smoothed L⁰ penalty to retain only essential parameters.
- The final model is extracted as a sparse, explicit functional form that is interpretable and directly usable in finite element solvers.
- The method is applied to hyperelasticity, yield functions, and isotropic hardening models using both synthetic and experimental data.
Experimental results
Research questions
- RQ1Can extreme sparsification of physics-augmented neural networks yield interpretable constitutive models without assuming a priori functional forms?
- RQ2How well can such a method generalize and extrapolate beyond training data, especially in low-data regimes?
- RQ3To what extent does enforcing physical constraints improve generalization and reduce overfitting in sparse, data-driven models?
- RQ4Can the method discover compact, human-readable functional forms for complex material behaviors like elastoplastic hardening?
- RQ5How does the smoothed L⁰-regularization compare to traditional pruning or sparse regression in terms of interpretability and accuracy?
Key findings
- The method successfully discovered interpretable constitutive models for compressible and incompressible hyperelasticity with high accuracy on synthetic data.
- For experimental SS316L stainless steel, the model achieved a fit with a hardening function of the form R(r) = 0.023 + 1.662/(1+1.071e^(-190.683r)) + 0.362/(1+1.071e^(-2200.640r)), showing no hardening at r=0.
- On 40Cr3MoV bainitic steel data, the model produced a complex yet interpretable hardening function with three active terms, achieving accurate interpolation and extrapolation.
- The loss and active parameter count decreased monotonically over training, indicating stable convergence and effective sparsity induction.
- The approach enabled reliable generalization beyond training strain ranges, even with limited data, due to physical constraint enforcement.
- The final sparse models were compact and analytically expressible, enabling direct integration into existing finite element frameworks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.