Skip to main content
QUICK REVIEW

[论文解读] Multimap targeted free energy estimation

Andrea Rizzi, Paolo Carloni|arXiv (Cornell University)|Feb 15, 2023
Machine Learning in Materials ScienceMaterials Science参考文献 94被引用 3
一句话总结

本文提出了一种多映射目标自由能微扰(TFEP)方法,通过使用来自廉价参考模拟的多个构型映射,加速了量子力学自由能计算,避免了昂贵的归一化流训练。该方法在对类药物分子从力场到DFTB3计算自由能差时,相比标准FEP实现了约3,000倍的速度提升,相比先前的非平衡方法也实现了约8倍的速度提升。

ABSTRACT

We present a new method to compute free energies at a quantum mechanical (QM) level of theory from molecular simulations using cheap reference potential energy functions, such as force fields. To overcome the poor overlap between the reference and target distributions, we generalize targeted free energy perturbation (TFEP) to employ multiple configuration maps. While TFEP maps have been obtained before from an expensive training of a normalizing flow neural network (NN), our multimap estimator allows us to use the same set of QM calculations to both optimize the maps and estimate the free energy, thus removing almost completely the overhead due to training. A multimap extension of the multistate Bennett acceptance ratio estimator is also derived for cases where samples from two or more states are available. Furthermore, we propose a one-epoch learning policy that can be used to efficiently avoid overfitting when computing the loss function is expensive compared to generating data. Finally, we show how our multimap approach can be combined with enhanced sampling strategies to overcome the pervasive problem of poor convergence due to slow degrees of freedom. We test our method on the HiPen dataset of drug-like molecules and fragments, and we show that it can accelerate the calculation of the free energy difference of switching from a force field to a DFTB3 potential by about 3 orders of magnitude compared to standard FEP and by a factor of about 8 compared to previously published nonequilibrium calculations.

研究动机与目标

  • 解决自由能计算中参考分布(如力场)与目标分布(如DFTB3)之间重叠性差的挑战。
  • 克服在绝热自由能计算中采样慢模式自由度带来的高计算成本。
  • 通过重用QM数据同时用于映射优化和自由能估计,消除昂贵的归一化流训练需求。
  • 利用增强采样和多态估计器结合多个映射,实现高效的自由能估计。
  • 开发一种单周期学习策略,以在损失函数评估成本高昂时防止过拟合。

提出的方法

  • 将目标自由能微扰(TFEP)推广至使用多个构型映射,而非单一映射。
  • 使用与自由能估计相同的QM数据训练多个映射,避免额外的训练开销。
  • 推导多态Bennett接受率(MBAR)估计器的多映射扩展,用于多态自由能估计。
  • 实施单周期学习策略,以在损失函数评估成本高昂时最小化过拟合。
  • 将多映射TFEP与增强采样技术(如OPES)结合,加速具有慢模式自由度系统的收敛。
  • 使用Z-矩阵或笛卡尔坐标作为映射的输入,通过批量处理提升计算效率。

实验结果

研究问题

  • RQ1当参考分布与目标分布重叠性差时,多个构型映射是否能提升自由能估计的准确性和效率?
  • RQ2通过重用QM数据同时用于映射学习和自由能估计,是否能消除归一化流的训练开销?
  • RQ3与标准FEP和非平衡方法相比,多映射TFEP方法在速度和准确性方面表现如何?
  • RQ4多映射方法是否能与增强采样技术有效结合,以加速具有慢模式自由度系统的收敛?
  • RQ5批量大小和坐标表示方式(笛卡尔坐标与Z-矩阵)对自由能估计的稳定性和准确性有何影响?

主要发现

  • 与标准绝热自由能微扰(FEP)相比,多映射TFEP方法在计算从力场到DFTB3的自由能差时,实现了约3,000倍的运行时间加速。
  • 与先前发表的非平衡FEP方法相比,该方法在HiPen数据集上的计算量减少了约8倍。
  • 多映射MBAR估计器能够利用来自多个状态的样本实现稳健的自由能估计,提升统计收敛性。
  • 使用Z-矩阵坐标并设置批量大小为48时,该方法在'良好'数据集上的平均绝对误差约为3.8 kcal/mol,置信区间为约0.04–0.08 kcal/mol。
  • 即使在损失函数评估成本高昂的情况下,单周期学习策略仍能有效防止过拟合,在10次重复模拟中保持稳定性能。
  • 对于收敛性较差的分子(如18号和19号),该方法仍提供不确定性较大的估计值(如18号分子为3.37 kcal/mol),表明采样质量存在局限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。