Skip to main content
QUICK REVIEW

[论文解读] Peering inside the black box: Learning the relevance of many-body functions in Neural Network potentials

Klara Bonneau, Jonas Lederer|arXiv (Cornell University)|Jul 5, 2024
Machine Learning in Materials Science被引用 7
一句话总结

本论文使用 GNN-LRP 来解释神经网络势在粗粒化分子系统中学习到的多体相互作用,展示在甲烷、水和 NTL9 蛋白质中的物理意义的 2-体和 3-体贡献。

ABSTRACT

Machine learned potentials are becoming a popular tool to define an effective energy model for complex systems, either incorporating electronic structure effects at the atomistic resolution, or effectively renormalizing part of the atomistic degrees of freedom at a coarse-grained resolution. One of the main criticisms to machine learned potentials is that the energy inferred by the network is not as interpretable as in more traditional approaches where a simpler functional form is used. Here we address this problem by extending tools recently proposed in the nascent field of Explainable Artificial Intelligence (XAI) to coarse-grained potentials based on graph neural networks (GNN). We demonstrate the approach on three different coarse-grained systems including two fluids (methane and water) and the protein NTL9. On these examples, we show that the neural network potentials can be in practice decomposed in relevance contributions to different orders, that can be directly interpreted and provide physical insights on the systems of interest.

研究动机与目标

  • 证明基于图神经网络(GNN)的粗粒化势能够编码可解释的多体相互作用。
  • 将层级相关传播(LRP)扩展到 GNN,以将网络能量分解为多体贡献。
  • 在简单流体上比较不同的 GNN 架构,以显示所学项的物理相关性的一致性。
  • 将该解释框架应用于蛋白质模型(NTL9),以识别稳定/不稳定相互作用并评估突变影响。

提出的方法

  • 使用力匹配以实现热力学一致性,从原子数据训练两个粗粒化 GNN 能量模型(PaiNN 和 SO3Net)。
  • 应用 GNN-LRP 基于在 GNN 中的遍历,将预测分解为 n-体相关分数(1- 到 (N_l+1)-体)。
  • 将 2-体与 3-体相关性与径向分布函数和角分布进行比较,以解释学习到的相互作用。
  • 分析 NTL9 CG 模型,将相关性映射到残基对和三元组,包含突变对相关性模式的影响。
Figure 1: Concept of GNN-LRP illustrated for a system of four particles (i.e. CG beads, in the present context). a) In GNNs, the input graph is defined by a cutoff radius that determines the direct neighbors for each input node. By stacking several message aggregations in multiple layers, informatio
Figure 1: Concept of GNN-LRP illustrated for a system of four particles (i.e. CG beads, in the present context). a) In GNNs, the input graph is defined by a cutoff radius that determines the direct neighbors for each input node. By stacking several message aggregations in multiple layers, informatio

实验结果

研究问题

  • RQ1基于 GNN 的粗粒化势是否可以分解为可解释的多体贡献,这些贡献与物理相互作用相对应?
  • RQ2对于同一系统,不同的 GNN 架构是否会产生一致的多体相关性模式?
  • RQ3在甲烷、水和 NTL9 中,哪些 2-体和 3-体相互作用对再现结构性质最为关键?
  • RQ4NTL9 的突变如何影响学习到的相关性以及结构模体的稳定性?

主要发现

  • GNN-LRP 将 CG 能量分解为有意义的 2-体和 3-体贡献,与物理直觉一致(如水中第一溶剂壳周围的稳定化)。
  • PaiNN 和 SO3Net 产生一致的定性多体相关性模式,尽管角度表示和截断半径不同,但两者都再现了关键的结构特征。
  • 3-体项对于再现水的角度结构至关重要;仅有 2-体的模型无法捕获水的正确角分布。
  • 在 NTL9 中,2-体相关性突出显示主要二级结构内的稳定接触,突变(ILE4ASN, LEU30PHE)扰动特定疏水/亲水相互作用,并在相关性模式的改变中体现出来。
  • 相关性映射区分折叠态、展开态和中间态,揭示与已知折叠模体一致的路径特异性相互作用。
Figure 2: Comparison of radial distribution functions resulting from simulations with an atomistic or CG model and corresponding 2-body relevance. Panels a) and c) correspond to water and b) and d) to methane models. Panels a) and b) show the results for PaiNN-based, and panels c) and d) for SO3Net-
Figure 2: Comparison of radial distribution functions resulting from simulations with an atomistic or CG model and corresponding 2-body relevance. Panels a) and c) correspond to water and b) and d) to methane models. Panels a) and b) show the results for PaiNN-based, and panels c) and d) for SO3Net-

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。