Skip to main content
QUICK REVIEW

[论文解读] Full-Atom Peptide Design with Geometric Latent Diffusion

Xiangzhe Kong, Jia, Yinjun|arXiv (Cornell University)|Feb 21, 2024
Machine Learning in Materials ScienceMaterials Science被引用 3
一句话总结

该论文提出PepGLAD,一种用于端到端全原子肽设计的几何潜在扩散模型,其条件基于目标结合位点。通过使用固定大小的潜在表征的变分自编码器(VAE)以及针对受体的仿射变换来处理全原子几何结构和可变的结合几何结构,实现了18%更高的多样性、8%更高的体外成功率,以及26%更高的参考结合构象召回率。

ABSTRACT

Peptide design plays a pivotal role in therapeutics, allowing brand new possibility to leverage target binding sites that are previously undruggable. Most existing methods are either inefficient or only concerned with the target-agnostic design of 1D sequences. In this paper, we propose a generative model for full-atom extbf{Pep}tide design with extbf{G}eometric extbf{LA}tent extbf{D}iffusion (PepGLAD) given the binding site. We first establish a benchmark consisting of both 1D sequences and 3D structures from Protein Data Bank (PDB) and literature for systematic evaluation. We then identify two major challenges of leveraging current diffusion-based models for peptide design: the full-atom geometry and the variable binding geometry. To tackle the first challenge, PepGLAD derives a variational autoencoder that first encodes full-atom residues of variable size into fixed-dimensional latent representations, and then decodes back to the residue space after conducting the diffusion process in the latent space. For the second issue, PepGLAD explores a receptor-specific affine transformation to convert the 3D coordinates into a shared standard space, enabling better generalization ability across different binding shapes. Experimental Results show that our method not only improves diversity and binding affinity significantly in the task of sequence-structure co-design, but also excels at recovering reference structures for binding conformation generation.

研究动机与目标

  • 实现基于目标结合位点的1D肽序列与3D全原子结构的端到端共同设计。
  • 通过VAE学习固定维度的潜在表征,解决扩散模型中可变大小全原子残基的挑战。
  • 通过受体特异性仿射变换将3D坐标映射到共享标准空间,提升在多样化结合几何结构上的泛化能力。
  • 构建一个结合PDB和文献数据的新训练数据集,用于全原子肽设计。
  • 与现有方法相比,实现更高的多样性、结合亲和力和构象准确性。

提出的方法

  • 训练变分自编码器(VAE)将可变大小的全原子残基编码为包含3D坐标和隐含特征的固定维度潜在向量,实现在一致潜在空间中的扩散。
  • VAE的编码器和解码器特别设计用于处理全原子输入和输出,在生成过程中保持原子级相互作用。
  • 从结合位点的中心偏移和协方差矩阵的Cholesky分解中推导出受体特异性仿射变换,将数据坐标映射到标准高斯空间。
  • 在共享标准空间中执行扩散过程,提升在多样化结合位点几何结构上的泛化能力,并支持对未见形状的迁移学习。
  • 通过将受体的3D结构同时作为VAE和扩散过程的条件,引导序列和结构生成以实现稳定且高亲和力的相互作用。
  • 该框架与Rosetta集成以进行能量评分和结合亲和力验证,使用PyRosetta计算dG_separated值以评估结合成功。

实验结果

研究问题

  • RQ1潜在扩散模型能否有效生成全原子肽结构,同时保持原子级相互作用?
  • RQ2如何在固定大小的扩散框架中对可变大小的全原子残基进行建模?
  • RQ3受体特异性仿射变换能否提升在多样化蛋白质-肽结合几何结构上的泛化能力?
  • RQ4与现有基线方法相比,该方法在多样性与体外结合成功率方面提升了多少?
  • RQ5该模型能否准确召回已知肽-受体复合物的参考结合构象?

主要发现

  • 与基线方法相比,PepGLAD在序列-结构共同设计中实现了18%的多样性提升。
  • 在序列-结构共同设计中,结合亲和力预测的体外成功率提高了8%。
  • 该模型在召回已知复合物的参考结合构象方面实现了26%的绝对提升。
  • 使用受体特异性仿射变换显著提升了在多样化结合位点几何结构上的泛化能力,即使对于未见过的形状也表现良好。
  • 基于VAE的潜在表征在扩散过程中成功保持了全原子几何结构,支持高保真度的结构生成。
  • 该模型在多样性与结合构象召回方面均优于强基线方法,包括RFDiffusion、FlexPepDock、AlphaFold2和HSRN。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。