Skip to main content
QUICK REVIEW

[Paper Review] Peering inside the black box: Learning the relevance of many-body functions in Neural Network potentials

Klara Bonneau, Jonas Lederer|arXiv (Cornell University)|Jul 5, 2024
Machine Learning in Materials Science7 citations
TL;DR

The paper uses GNN-LRP to interpret the multi-body interactions learned by neural-network potentials for coarse-grained molecular systems, illustrating physically meaningful 2-body and 3-body contributions in methane, water, and the NTL9 protein.

ABSTRACT

Machine learned potentials are becoming a popular tool to define an effective energy model for complex systems, either incorporating electronic structure effects at the atomistic resolution, or effectively renormalizing part of the atomistic degrees of freedom at a coarse-grained resolution. One of the main criticisms to machine learned potentials is that the energy inferred by the network is not as interpretable as in more traditional approaches where a simpler functional form is used. Here we address this problem by extending tools recently proposed in the nascent field of Explainable Artificial Intelligence (XAI) to coarse-grained potentials based on graph neural networks (GNN). We demonstrate the approach on three different coarse-grained systems including two fluids (methane and water) and the protein NTL9. On these examples, we show that the neural network potentials can be in practice decomposed in relevance contributions to different orders, that can be directly interpreted and provide physical insights on the systems of interest.

Motivation & Objective

  • Demonstrate that graph neural network (GNN) based coarse-grained potentials encode interpretable many-body interactions.
  • Extend layerwise-relevance propagation (LRP) to GNNs to decompose network energy into multi-body contributions.
  • Compare different GNN architectures on simple fluids to show consistent physical relevance of learned terms.
  • Apply the interpretation framework to a protein model (NTL9) to identify stabilizing/destabilizing interactions and assess mutations.

Proposed method

  • Train two coarse-grained GNN energy models (PaiNN and SO3Net) from atomistic data using force-matching for thermodynamic consistency.
  • Apply GNN-LRP to decompose predictions into n-body relevance scores (1- to (N_l+1)-body) based on walks in the GNN.
  • Compare 2-body and 3-body relevance with radial distribution functions and angular distributions to interpret learned interactions.
  • Analyze NTL9 CG model to map relevance onto residue pairs and triplets, including effects of mutations on relevance patterns.
Figure 1: Concept of GNN-LRP illustrated for a system of four particles (i.e. CG beads, in the present context). a) In GNNs, the input graph is defined by a cutoff radius that determines the direct neighbors for each input node. By stacking several message aggregations in multiple layers, informatio
Figure 1: Concept of GNN-LRP illustrated for a system of four particles (i.e. CG beads, in the present context). a) In GNNs, the input graph is defined by a cutoff radius that determines the direct neighbors for each input node. By stacking several message aggregations in multiple layers, informatio

Experimental results

Research questions

  • RQ1Can GNN-based coarse-grained potentials be decomposed into interpretable multi-body contributions that correspond to physical interactions?
  • RQ2Do different GNN architectures yield consistent multi-body relevance patterns for the same systems?
  • RQ3What 2-body and 3-body interactions are most responsible for reproducing structural properties in methane, water, and NTL9?
  • RQ4How do mutations in NTL9 affect the learned relevance and the stability of structural motifs?

Key findings

  • GNN-LRP decomposes the CG energy into meaningful 2-body and 3-body contributions that align with physical intuition (e.g., stabilization around first solvation shell in water).
  • PaiNN and SO3Net yield consistent qualitative multi-body relevance patterns, though angular representations and cutoffs differ, yet both reproduce key structural features.
  • 3-body terms are crucial for reproducing water’s angular structure; a 2-body only model fails to capture the correct angular distributions for water.
  • In NTL9, 2-body relevance highlights stabilizing contacts within major secondary structures, and mutations (ILE4ASN, LEU30PHE) perturb specific hydrophobic/hydrophilic interactions as reflected in altered relevance patterns.
  • Relevance maps differentiate folded vs. unfolded and intermediate states, revealing pathway-specific interactions consistent with known folding motifs.
Figure 2: Comparison of radial distribution functions resulting from simulations with an atomistic or CG model and corresponding 2-body relevance. Panels a) and c) correspond to water and b) and d) to methane models. Panels a) and b) show the results for PaiNN-based, and panels c) and d) for SO3Net-
Figure 2: Comparison of radial distribution functions resulting from simulations with an atomistic or CG model and corresponding 2-body relevance. Panels a) and c) correspond to water and b) and d) to methane models. Panels a) and b) show the results for PaiNN-based, and panels c) and d) for SO3Net-

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.