Skip to main content
QUICK REVIEW

[Paper Review] Protein Language Models and Structure Prediction: Connection and Progression

Bozhen Hu, Jun Xia|arXiv (Cornell University)|Nov 30, 2022
Machine Learning in Bioinformatics21 citations
TL;DR

A systematic survey linking protein language models (pLMs) to protein structure prediction (PSP), detailing architectures, pre-training, databases, and applications, with future directions.

ABSTRACT

The prediction of protein structures from sequences is an important task for function prediction, drug design, and related biological processes understanding. Recent advances have proved the power of language models (LMs) in processing the protein sequence databases, which inherit the advantages of attention networks and capture useful information in learning representations for proteins. The past two years have witnessed remarkable success in tertiary protein structure prediction (PSP), including evolution-based and single-sequence-based PSP. It seems that instead of using energy-based models and sampling procedures, protein language model (pLM)-based pipelines have emerged as mainstream paradigms in PSP. Despite the fruitful progress, the PSP community needs a systematic and up-to-date survey to help bridge the gap between LMs in the natural language processing (NLP) and PSP domains and introduce their methodologies, advancements and practical applications. To this end, in this paper, we first introduce the similarities between protein and human languages that allow LMs extended to pLMs, and applied to protein databases. Then, we systematically review recent advances in LMs and pLMs from the perspectives of network architectures, pre-training strategies, applications, and commonly-used protein databases. Next, different types of methods for PSP are discussed, particularly how the pLM-based architectures function in the process of protein folding. Finally, we identify challenges faced by the PSP community and foresee promising research directions along with the advances of pLMs. This survey aims to be a hands-on guide for researchers to understand PSP methods, develop pLMs and tackle challenging problems in this field for practical purposes.

Motivation & Objective

  • Explain how protein sequences are treated like natural language and why LMs are applicable to PSP.
  • Survey network architectures, pre-training strategies, and applications of pLMs in PSP.
  • Summarize commonly used protein databases and their role in pre-training.
  • Discuss how pLM-based architectures integrate into PSP pipelines and folding mechanisms.
  • Identify challenges and propose promising directions for future research in pLMs and PSP.

Proposed method

  • Review and categorize existing protein language models (pLMs) and their pre-training methods.
  • Analyze how pLMs are incorporated into PSP pipelines and structure-feature learning.
  • Compare evolution-based and single-sequence-based PSP approaches in the context of pLMs.
  • Summarize protein databases and data resources used for pre-training pLMs.
  • Highlight limitations and forecast future trends in the field.

Experimental results

Research questions

  • RQ1How do pLMs relate to traditional PSP methods in terms of data, representations, and objectives?
  • RQ2What architectural choices and pre-training strategies are most effective for PSP tasks?
  • RQ3What are the key databases and data resources underpinning pLM development for PSP?
  • RQ4What are the current limitations of pLM-based PSP methods and what future directions are promising?

Key findings

  • pLMs leverage large-scale protein sequence data to learn representations useful for structure and function prediction.
  • Transformer-based pLMs, combined with pre-training on diverse databases, have driven progress in PSP, including single-sequence and evolution-based approaches.
  • Various pre-training strategies (e.g., masked language modeling) and architectures (RNN/LSTM vs. Transformer) are employed to capture protein dependencies.
  • pLMs enable PSP pipelines that can operate with limited or no evolutionary information in some contexts.
  • The survey synthesizes methods for antibody, protein complex, protein–ligand, and protein–RNA structure-related tasks.
  • Challenges include tokenization of protein sequences, limited labeled data, and the need for integrated multi-modal data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.