Skip to main content
QUICK REVIEW

[论文解读] Generative artificial intelligence for de novo protein design

Adam Winnifrith, Carlos Outeiral|arXiv (Cornell University)|Oct 15, 2023
Chemical Synthesis and Analysis被引用 5
一句话总结

本文综述了生成式人工智能方法在从头蛋白设计中的应用,重点探讨了语言模型与扩散过程如何实现新颖且具有功能性的蛋白质设计,其实验成功率接近20%。文章强调通过整合生物化学知识,提升模型在设计具有复杂行为(如构象变化和翻译后调控)蛋白质时的性能与可解释性。

ABSTRACT

Engineering new molecules with desirable functions and properties has the potential to extend our ability to engineer proteins beyond what nature has so far evolved. Advances in the so-called "de novo" design problem have recently been brought forward by developments in artificial intelligence. Generative architectures, such as language models and diffusion processes, seem adept at generating novel, yet realistic proteins that display desirable properties and perform specified functions. State-of-the-art design protocols now achieve experimental success rates nearing 20%, thus widening the access to de novo designed proteins. Despite extensive progress, there are clear field-wide challenges, for example in determining the best in silico metrics to prioritise designs for experimental testing, and in designing proteins that can undergo large conformational changes or be regulated by post-translational modifications and other cellular processes. With an increase in the number of models being developed, this review provides a framework to understand how these tools fit into the overall process of de novo protein design. Throughout, we highlight the power of incorporating biochemical knowledge to improve performance and interpretability.

研究动机与目标

  • 提供一个全面的框架,以理解生成式人工智能在从头蛋白设计中在整个设计流程中的作用。
  • 识别并解决在选择计算指标以优先筛选实验测试设计时的关键挑战。
  • 探索能够实现大规模构象变化及通过翻译后修饰实现调控的蛋白质设计。
  • 通过将特定领域的生物化学知识整合到生成架构中,提升模型性能与可解释性。
  • 指导研究人员选择并应用最有效的生成模型以实现实验性蛋白工程。

提出的方法

  • 利用基于Transformer的语言模型,基于蛋白质序列数据进行训练,以生成新颖且进化上合理的氨基酸序列。
  • 应用基于扩散的生成模型,通过迭代去噪潜在表征,生成具有高真实感的3D蛋白质结构。
  • 将生物化学先验知识(如残基稳定性、折叠倾向性及功能基序)整合到生成模型的潜在空间中。
  • 采用多阶段设计流程,结合条件生成与结构感知损失函数,确保设计的结构与功能合理性。
  • 利用计算指标(如预测的ΔΔG值、结构质量指标如RMSD、功能位点保守性)评估模型输出。
  • 利用在精选蛋白质数据库上的迁移学习与微调,以提升泛化能力并减少分布偏移。

实验结果

研究问题

  • RQ1如何有效构建生成式AI模型,以实现高实验成功率的新颖功能性蛋白序列设计?
  • RQ2哪些计算指标最能预测从头蛋白设计中的实验成功率?这些指标在生成过程中如何优化?
  • RQ3在多大程度上可通过整合生物化学知识来增强生成模型,以提升其结构与功能保真度?
  • RQ4生成式模型能否可靠地设计出可发生大规模构象变化或受翻译后修饰调控的蛋白质?
  • RQ5在从头蛋白设计中,自回归Transformer与扩散模型等不同生成架构在性能与可解释性方面如何比较?

主要发现

  • 最先进的生成模型现已实现接近20%的从头蛋白设计实验成功率,显著提升了实验验证的可行性。
  • 将生物化学先验知识整合到生成模型的潜在空间中,可同时提升生成蛋白的结构合理性与功能相关性。
  • 基于扩散的模型在生成具有高几何精度、RMSD接近天然折叠结构的多样化且稳定的3D蛋白结构方面表现优异。
  • 基于语言模型的方法在生成序列优化的蛋白质方面表现出色,可实现期望的功能基序与稳定性特征。
  • 整合如预测ΔΔG值和结构质量评分等计算指标,可有效提升设计优先级,便于实验筛选。
  • 尽管已取得进展,但在设计具有动态行为(如别构效应或受调控的折叠)的蛋白质方面仍面临挑战,表明需对细胞过程进行更复杂的建模。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。