Skip to main content
QUICK REVIEW

[论文解读] PepMLM: Target Sequence-Conditioned Generation of Therapeutic Peptide Binders via Span Masked Language Modeling

Tianlai Chen, Dumas, Madeleine|PubMed|Oct 5, 2023
Protein Degradation and Inhibitors参考文献 29被引用 17
一句话总结

PepMLM 是一个以目标序列为条件的生成器,利用 span 掩码策略设计全新肽结合体,并对蛋白质语言模型进行微调,同时对结合体效力进行计算机与实验验证。

ABSTRACT

Target proteins that lack accessible binding pockets and conformational stability have posed increasing challenges for drug development. Induced proximity strategies, such as PROTACs and molecular glues, have thus gained attention as pharmacological alternatives, but still require small molecule docking at binding pockets for targeted protein degradation. The computational design of protein-based binders presents unique opportunities to access "undruggable" targets, but have often relied on stable 3D structures or structure-influenced latent spaces for effective binder generation. In this work, we introduce <b>PepMLM</b>, a target sequence-conditioned generator of <i>de novo</i> linear peptide binders. By employing a novel span masking strategy that uniquely positions cognate peptide sequences at the C-terminus of target protein sequences, PepMLM fine-tunes the state-of-the-art ESM-2 pLM to fully reconstruct the binder region, achieving low perplexities matching or improving upon validated peptide-protein sequence pairs. After successful <i>in silico</i> benchmarking with AlphaFold-Multimer, outperforming RFDiffusion on structured targets, we experimentally verify PepMLM's efficacy via fusion of model-derived peptides to E3 ubiquitin ligase domains, demonstrating endogenous degradation of emergent viral phosphoproteins and Huntington's disease-driving proteins. In total, PepMLM enables the generative design of candidate binders to any target protein, without the requirement of target structure, empowering downstream therapeutic applications.

研究动机与目标

  • 解决如何为缺乏可访问结合口袋或稳定结构的靶标生成 de novo 肽结合体。
  • 开发一个以靶标序列为条件的生成器,将同源结合体序列放置在 C 端以实现有效重建。
  • 使用 AlphaFold-Multimer 基准测试和实验降解测定来评估该方法。
  • 实现对任何靶标蛋白的结合体设计,而不依赖靶标结构。

提出的方法

  • 引入将认知结合体序列定位在靶标 C 端的 span 掩码。
  • 对最先进的 ESM-2 蛋白质语言模型进行微调,以重建结合区。
  • 在结构化靶标上与 AlphaFold-Multimer 进行计算机基准测试并与 RFDiffusion 进行比较。
  • 通过将生成的结合体与 E3 疏变连接域融合以实现靶标蛋白的降解来进行实验验证。
  • 评估在不需要靶标结构知识的情况下生成候选结合体的能力。

实验结果

研究问题

  • RQ1基于靶标序列条件的 span 掩码策略是否能实现对结合区域的可靠生成?
  • RQ2在此条件下对蛋白质语言模型进行微调是否能提升结合体重建的可信度与合理性?
  • RQ3在结构化靶标上,PepMLM 相较于现有扩散法在计算机评估中的表现如何?
  • RQ4PepMLM 派生的结合体在实验测定中是否能有效驱动靶标蛋白的降解?

主要发现

  • PepMLM 在结合体重建上实现了与经过验证的结合对相比不低的困惑度,甚至有所提升。
  • 使用 AlphaFold-Multimer 的计算基准测试支持 PepMLM 的有效性,并在结构化靶标上优于 RFDiffusion。
  • 实验验证显示模型派生肽在融合到 E3 泛素连接域后可诱导内源性病毒磷酸化蛋白和亨廷顿病驱动蛋白的降解。
  • 该方法能够在不要求靶标结构知识的情况下为任意靶标蛋白生成候选结合体。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。