Skip to main content
QUICK REVIEW

[Paper Review] Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling

Ahmed Elnaggar, Hazem Essam|arXiv (Cornell University)|Jan 16, 2023
Machine Learning in BioinformaticsBiochemistry, Genetics and Molecular Biology27 citations
TL;DR

Ankh proposes protein-specific optimization for language models, achieving general-purpose modeling with substantially smaller pre-training data, inference size, and embedding dimensions while surpassing state-of-the-art on protein benchmarks.

ABSTRACT

As opposed to scaling-up protein language models (PLMs), we seek improving performance via protein-specific optimization. Although the proportionality between the language model size and the richness of its learned representations is validated, we prioritize accessibility and pursue a path of data-efficient, cost-reduced, and knowledge-guided optimization. Through over twenty experiments ranging from masking, architecture, and pre-training data, we derive insights from protein-specific experimentation into building a model that interprets the language of life, optimally. We present Ankh, the first general-purpose PLM trained on Google's TPU-v4 surpassing the state-of-the-art performance with fewer parameters (<10% for pre-training, <7% for inference, and <30% for the embedding dimension). We provide a representative range of structure and function benchmarks where Ankh excels. We further provide a protein variant generation analysis on High-N and One-N input data scales where Ankh succeeds in learning protein evolutionary conservation-mutation trends and introducing functional diversity while retaining key structural-functional characteristics. We dedicate our work to promoting accessibility to research innovation via attainable resources.

Motivation & Objective

  • Improve protein language model performance through data-efficient, cost-reduced, and knowledge-guided optimization rather than scaling up model size.
  • Investigate the impact of masking, architecture, and pre-training data choices to derive protein-specific insights for general-purpose modelling.
  • Demonstrate that a smaller, optimized model can surpass state-of-the-art on diverse structure and function benchmarks.
  • Analyze protein variant generation under High-N and One-N data scales to learn evolutionary conservation-mutation trends and functional diversity.
  • Promote accessibility by providing attainable resources and an open pathway to research innovation.

Proposed method

  • Experiment with over twenty protein-specific design choices spanning masking, architecture, and pre-training data.
  • Train Ankh, a general-purpose PLM, on Google's TPU-v4 hardware.
  • Compare against state-of-the-art PLMs using a representative set of structure and function benchmarks.
  • Evaluate protein variant generation under High-N and One-N input data scales to assess conservation, mutation trends, and functional diversity.
  • Analyze how fewer parameters and reduced embedding dimension affect performance and accessibility.

Experimental results

Research questions

  • RQ1Can protein-specific optimization yield general-purpose PLM performance competitive with or surpassing larger models without scaling up?
  • RQ2What masking, architectural, and data choices most improve protein language understanding and downstream tasks?
  • RQ3How does Ankh perform on structure/function benchmarks compared to prior state-of-the-art PLMs?
  • RQ4Does Ankh learn evolutionary conservation-mutation trends and support functional diversity under constrained data scales?
  • RQ5What are the resource implications (pre-training data, inference, embedding size) of effective protein-specific PLMs?

Key findings

  • Ankh surpasses state-of-the-art performance with fewer parameters and substantially reduced resources.
  • Pre-training requires <10% of the parameters, inference uses <7% of the parameters, and embedding dimension is <30% of a typical baseline.
  • Ankh shows strong performance across a representative range of structure and function benchmarks.
  • Under High-N and One-N data scales, Ankh learns evolutionary conservation-mutation trends and introduces functional diversity while preserving key structural-functional characteristics.
  • The work emphasizes accessibility by prioritizing data-efficient optimization and attainable resources.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.