Skip to main content
QUICK REVIEW

[论文解读] Interpreting BERT architecture predictions for peptide presentation by MHC class I proteins

Hans-Christof Gasser, Georges Bedran|arXiv (Cornell University)|Nov 13, 2021
vaccines and immunoinformatics approaches被引用 6
一句话总结

本文介绍了 ImmunoBERT,一种基于 BERT 的深度学习模型,通过肽序列及其侧翼区域与 MHC-I 等位基因,预测 MHC I 类分子对肽的呈递。通过应用 SHAP 和 LIME 可解释性方法,作者识别出肽的 N- 和 C-末端残基以及特定 MHC-I 沉积袋残基(A、B、F)在呈递过程中起关键作用,结合 3D 可视化和基序分析,验证了其生物学机制。

ABSTRACT

The major histocompatibility complex (MHC) class-I pathway supports the detection of cancer and viruses by the immune system. It presents parts of proteins (peptides) from inside a cell on its membrane surface enabling visiting immune cells that detect non-self peptides to terminate the cell. The ability to predict whether a peptide will get presented on MHC Class I molecules helps in designing vaccines so they can activate the immune system to destroy the invading disease protein. We designed a prediction model using a BERT-based architecture (ImmunoBERT) that takes as input a peptide and its surrounding regions (N and C-terminals) along with a set of MHC class I (MHC-I) molecules. We present a novel application of well known interpretability techniques, SHAP and LIME, to this domain and we use these results along with 3D structure visualizations and amino acid frequencies to understand and identify the most influential parts of the input amino acid sequences contributing to the output. In particular, we find that amino acids close to the peptides' N- and C-terminals are highly relevant. Additionally, some positions within the MHC proteins (in particular in the A, B and F pockets) are often assigned a high importance ranking - which confirms biological studies and the distances in the structure visualizations.

研究动机与目标

  • 开发一种深度学习模型,以准确预测 MHC I 类分子对肽的呈递,这是疫苗设计中的关键步骤。
  • 利用与模型无关的可解释性技术(SHAP 和 LIME)解释模型预测结果,识别出具有生物学意义的序列特征。
  • 通过肽-MHC 相互作用的 3D 结构可视化和序列基序分析,验证模型的解释结果。
  • 展示模型在多种 MHC-I 等位基因(包括罕见等位基因)上的泛化能力,以支持个性化免疫治疗。
  • 通过使深度学习预测结果对生物学家和临床医生可解释,弥合计算与生物学理解之间的鸿沟。

提出的方法

  • 该模型 ImmunoBERT 采用基于 BERT 的架构,处理包含 9-肽、其 N-侧翼和 C-侧翼区域以及 MHC-I 等位基因序列的输入序列。
  • 模型在实验测定的肽-MHC-I 结合亲和力数据集上进行微调,以预测呈递可能性。
  • 应用 SHAP 和 LIME 生成肽、侧翼区域和 MHC-I 蛋白中每个氨基酸的特征重要性分数。
  • 利用肽-MHC 复合物的 3D 结构可视化,验证 SHAP 和 LIME 识别出的高重要性残基在空间上的相关性。
  • 对预测的免疫原性肽进行序列基序分析,识别出与已知生物学基序一致的重复模式。
  • 将模型性能与 SOTA 方法(如 NetMHCpan 和 MHCflurry)进行基准对比,结果表明其预测准确性相当。

实验结果

研究问题

  • RQ1肽和 MHC-I 分子的哪些区域在决定 MHC-I 对肽的呈递中最具影响力?
  • RQ2SHAP 和 LIME 的特征重要性分数在多大程度上与已知的 MHC-I 抗原呈递的结构和生物学机制一致?
  • RQ3模型的预测和解释结果在不同 MHC-I 等位基因(包括罕见等位基因)上如何泛化?
  • RQ43D 结构可视化和序列基序分析是否能佐证 SHAP 和 LIME 的可解释性发现?
  • RQ5肽侧翼区域(N-和 C-侧翼)对呈递预测的贡献是什么?其与核心肽的贡献相比如何?

主要发现

  • N-和 C-末端附近的肽残基在 SHAP 和 LIME 中均被一致赋予高重要性,表明其在 MHC-I 呈递中起关键作用。
  • MHC-I 分子中 A、B 和 F 沉积袋的特定残基——特别是 HLA-B*54:01 的 66、95 和 116 位——被识别为极具影响力,与已知的肽结合结构生物学一致。
  • 模型对 HLA-B*54:01 的预测基序显示,在肽的第 2 和第 9 位富集疏水性残基(如丙氨酸、缬氨酸、亮氨酸),与已知的结合偏好一致。
  • 3D 可视化证实,高重要性 MHC 残基(如 95 位的色氨酸和 116 位的亮氨酸)分别与肽的 C-末端和 F-口袋形成疏水相互作用。
  • 肽侧翼区域对预测的贡献小于核心肽,但当包含时,仍能带来小而一致的性能提升。
  • 模型在未见过的 MHC-I 等位基因上表现出泛化能力,表现为在保留等位基因上保持一致的基序模式和可解释性结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。