Skip to main content
QUICK REVIEW

[论文解读] A 5' UTR Language Model for Decoding Untranslated Regions of mRNA and Function Predictions

Yanyi Chu, Dan Yu|arXiv (Cornell University)|Oct 5, 2023
RNA and protein synthesis mechanisms被引用 5
一句话总结

该论文提出UTR-LM,一种在多物种5' UTR上预训练的深度学习语言模型,并通过二级结构和最小自由能等结构特征进行微调。该模型在预测翻译效率和mRNA表达水平方面提升最高达60%,并成功实现了对211个合成5' UTR的实验验证,使蛋白产量提高32.5%。

ABSTRACT

The 5' UTR, a regulatory region at the beginning of an mRNA molecule, plays a crucial role in regulating the translation process and impacts the protein expression level. Language models have showcased their effectiveness in decoding the functions of protein and genome sequences. Here, we introduced a language model for 5' UTR, which we refer to as the UTR-LM. The UTR-LM is pre-trained on endogenous 5' UTRs from multiple species and is further augmented with supervised information including secondary structure and minimum free energy. We fine-tuned the UTR-LM in a variety of downstream tasks. The model outperformed the best-known benchmark by up to 42% for predicting the Mean Ribosome Loading, and by up to 60% for predicting the Translation Efficiency and the mRNA Expression Level. The model also applies to identifying unannotated Internal Ribosome Entry Sites within the untranslated region and improves the AUPR from 0.37 to 0.52 compared to the best baseline. Further, we designed a library of 211 novel 5' UTRs with high predicted values of translation efficiency and evaluated them via a wet-lab assay. Experiment results confirmed that our top designs achieved a 32.5% increase in protein production level relative to well-established 5' UTR optimized for therapeutics.

研究动机与目标

  • 开发一种深度学习模型,以解码5' UTR中调控功能,这些功能对翻译控制和蛋白质表达至关重要。
  • 提高对关键翻译调控特征(如核糖体装载、翻译效率和mRNA表达水平)的预测准确性。
  • 利用监督结构数据,在5' UTR中识别未注释的内部核糖体进入位点(IRES)。
  • 设计并实验验证具有优化翻译效率的合成5' UTR,以支持生物技术应用。

提出的方法

  • 在多个物种的内源性5' UTR序列上预训练UTR-LM,以学习序列水平的调控模式。
  • 通过二级结构和最小自由能的监督数据增强预训练模型,以提升对结构上下文的理解。
  • 在下游任务(包括核糖体装载、翻译效率和mRNA表达水平预测)上对UTR-LM进行微调。
  • 利用模型筛选潜在的内部核糖体进入位点(IRES),通过识别具有高结构稳定性的IRES样序列基序。
  • 基于UTR-LM输出,设计一个包含211个序列的合成5' UTR文库,预测其具有高翻译效率。
  • 通过湿实验检测验证表现最佳的合成5' UTR,以测量实际的蛋白产量水平。

实验结果

研究问题

  • RQ1在多物种5' UTR序列上预训练的深度学习语言模型,是否能在预测翻译效率和核糖体装载方面超越现有基准?
  • RQ2整合二级结构和最小自由能数据在多大程度上提升了5' UTR语言模型的预测性能?
  • RQ3该模型能否以高于当前最先进方法的精度识别未注释的内部核糖体进入位点(IRES)?
  • RQ4基于UTR-LM指导设计的合成5' UTR是否能在实验验证中实现可测量的蛋白表达提升?

主要发现

  • UTR-LM在预测平均核糖体装载方面相比最佳现有基准最高提升42%。
  • 与最佳基线相比,该模型在翻译效率和mRNA表达水平预测准确性上最高提升60%。
  • UTR-LM提升了IRES检测性能,AUPR从0.37提升至0.52,优于最强基线。
  • 合成文库中表现最佳的5' UTR设计使蛋白产量相比一个广泛认可的治疗性5' UTR提高了32.5%。
  • 由于在内源性5' UTR上进行多物种预训练,该模型在跨物种间表现出强大的泛化能力。
  • 整合结构特征(二级结构和最小自由能)显著提升了下游预测性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。