Skip to main content

Martin Stiege

Seoul National University · 生化学・遺伝学・分子生物学

研究室紹介

Professor Martin Stiege's research lab specializes in computational biology and bioinformatics, focusing on accelerating protein structure prediction and functional annotation through advanced sequence clustering and AI-driven methods. The lab develops high-performance software tools such as MMseqs2, ColabFold, and Linclust to enable large-scale analysis of genomic and metagenomic data with unprecedented speed and sensitivity. Their work centers on creating scalable, open-source solutions for protein sequence and structure prediction, particularly for challenging environmental and uncultivated microbial sequences. The lab also pioneers novel databases like Uniclust and Uniboost to enhance the accuracy and consistency of functional annotation across diverse proteomes.

protein structure predictionsequence clusteringmetagenomicsAI in biologybioinformatics tools

Research Overview

Papers
175
Total Citations
78,854
Papers (5y)
111
Primary Field
生化学・遺伝学・分子生物学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
111total
2022
2023
2024
2025
2026
Citations per year (5y)
17,484total
20222023202420252026

Selected Papers

15
1
Article|9,675 citations·2022
ColabFold: making protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, Martin Steinegger
SJR Q1Nature MethodsOA

ColabFold offers accelerated prediction of protein structures and complexes by combining the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. ColabFold's 40-60-fold faster search and optimized model utilization enables prediction of close to 1,000 structures per day on a server with one graphics processing unit. Coupled with Google Colaboratory, ColabFold becomes a free and accessible platform for protein folding. ColabFold is open-source software available at https://github.com/s

Molecular BiologyBiochemistry, Genetics and Molecular Biology
2
Article|5,115 citations·2017
MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets
Martin Steinegger, Johannes Söding
SJR Q1Nature Biotechnology
Molecular BiologyBiochemistry, Genetics and Molecular Biology
3
Article|2,373 citations·2023
Fast and accurate protein structure search with Foldseek
Michel van Kempen, Stephanie Kim, Charlotte Tumescheit, Milot Mirdita, Jeong-Jae Lee, Cameron L. M. Gilchrist, Johannes Söding, Martin Steinegger
SJR Q1Nature BiotechnologyOA

As structure prediction methods are generating millions of publicly available protein structures, searching these databases is becoming a bottleneck. Foldseek aligns the structure of a query protein against a database by describing tertiary amino acid interactions within proteins as sequences over a structural alphabet. Foldseek decreases computation times by four to five orders of magnitude with 86%, 88% and 133% of the sensitivities of Dali, TM-align and CE, respectively.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
4
Article|1,274 citations·2019
HH-suite3 for fast remote homology detection and deep protein annotation
Martin Steinegger, Markus Meier, Milot Mirdita, Harald Vöhringer, Stephan J. Haunsberger, Johannes Söding
SJR Q1BMC BioinformaticsOA

BACKGROUND: HH-suite is a widely used open source software suite for sensitive sequence similarity searches and protein fold recognition. It is based on pairwise alignment of profile Hidden Markov models (HMMs), which represent multiple sequence alignments of homologous proteins. RESULTS: We developed a single-instruction multiple-data (SIMD) vectorized implementation of the Viterbi algorithm for profile HMM alignment and introduced various other speed-ups. These accelerated the search methods H

Molecular BiologyBiochemistry, Genetics and Molecular Biology
5
Article|989 citations·2018
Clustering huge protein sequence sets in linear time
Martin Steinegger, Johannes Söding
SJR Q1Nature CommunicationsOA

Metagenomic datasets contain billions of protein sequences that could greatly enhance large-scale functional annotation and structure prediction. Utilizing this enormous resource would require reducing its redundancy by similarity clustering. However, clustering hundreds of millions of sequences is impractical using current algorithms because their runtimes scale as the input set size N times the number of clusters K, which is typically of similar order as N, resulting in runtimes that increase

Molecular BiologyBiochemistry, Genetics and Molecular Biology
6
Article|903 citations·2016
Uniclust databases of clustered and deeply annotated protein sequences and alignments
Milot Mirdita, Lars von den Driesch, Clovis Galiez, María Martin, Johannes Söding, Martin Steinegger
SJR Q1Nucleic Acids ResearchOA

We present three clustered protein sequence databases, Uniclust90, Uniclust50, Uniclust30 and three databases of multiple sequence alignments (MSAs), Uniboost10, Uniboost20 and Uniboost30, as a resource for protein sequence analysis, function prediction and sequence searches. The Uniclust databases cluster UniProtKB sequences at the level of 90%, 50% and 30% pairwise sequence identity. Uniclust90 and Uniclust50 clusters showed better consistency of functional annotation than those of UniRef90 an

Molecular BiologyBiochemistry, Genetics and Molecular Biology
7
Review|734 citations·2022
Metagenome analysis using the Kraken software suite
Jennifer Lu, Natalia Rincon, Derrick E. Wood, Florian P. Breitwieser, Christopher Pockrandt, Ben Langmead, Steven L. Salzberg, Martin Steinegger
SJR Q1Nature ProtocolsOA
Molecular BiologyBiochemistry, Genetics and Molecular Biology
8
Article|689 citations·2019
MMseqs2 desktop and local web server app for fast, interactive sequence searches
Milot Mirdita, Martin Steinegger, Johannes Söding
SJR Q1BioinformaticsOA

SUMMARY: The MMseqs2 desktop and web server app facilitates interactive sequence searches through custom protein sequence and profile databases on personal workstations. By eliminating MMseqs2's runtime overhead, we reduced response times to a few seconds at sensitivities close to BLAST. AVAILABILITY AND IMPLEMENTATION: The app is easy to install for non-experts. GPLv3-licensed code, pre-built desktop app packages for Windows, MacOS and Linux, Docker images for the web server application and a d

Molecular BiologyBiochemistry, Genetics and Molecular Biology
9
Preprint|589 citations·2021
ColabFold - Making protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, Martin Steinegger
bioRxiv (Cold Spring Harbor Laboratory)OA

ColabFold offers accelerated protein structure and complex predictions by combining the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. ColabFold’s 40 - 60× faster search and optimized model use allows predicting close to a thousand structures per day on a server with one GPU. Coupled with Google Colaboratory, ColabFold becomes a free and accessible platform for protein folding. ColabFold is open-source software available at github.com/sokrypton/ColabFold . Its novel environmenta

Molecular BiologyBiochemistry, Genetics and Molecular Biology
10
Article|471 citations·2019
Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold
Martin Steinegger, Milot Mirdita, Johannes Söding
SJR Q1Nature MethodsOA

The open-source de novo protein-level assembler, Plass ( https://plass.mmseqs.com ), assembles six-frame-translated sequencing reads into protein sequences. It recovers 2-10 times more protein sequences from complex metagenomes and can assemble huge datasets. We assembled two redundancy-filtered reference protein catalogs, 2 billion sequences from 640 soil samples (soil reference protein catalog) and 292 million sequences from 775 marine eukaryotic metatranscriptomes (marine eukaryotic reference

Molecular BiologyBiochemistry, Genetics and Molecular Biology
11
Preprint|378 citations·2022
Fast and accurate protein structure search with Foldseek
Michel van Kempen, Stephanie Kim, Charlotte Tumescheit, Milot Mirdita, Jeong-Jae Lee, Cameron L. M. Gilchrist, Johannes Söding, Martin Steinegger
bioRxiv (Cold Spring Harbor Laboratory)OA

As structure prediction methods are generating millions of publicly available protein structures, searching these databases is becoming a bottleneck. Foldseek aligns the structure of a query protein against a database by describing the amino acid backbone of proteins as sequences over a structural alphabet. Foldseek decreases computation times by four to five orders of magnitude with 86%, 88% and 133% of the sensitivities of DALI, TM-align and CE, respectively.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
12
Article|365 citations·2023
Clustering predicted structures at the scale of the known protein universe
Inigo Barrio‐Hernandez, Jingi Yeo, Jürgen Jänes, Milot Mirdita, Cameron L. M. Gilchrist, Tanita Wein, Mihály Váradi, Sameer Velankar, Pedro Beltrão, Martin Steinegger
SJR Q1NatureOA

Abstract Proteins are key to all cellular processes and their structure is important in understanding their function and evolution. Sequence-based predictions of protein structures have increased in accuracy 1 , and over 214 million predicted structures are available in the AlphaFold database 2 . However, studying protein structures at this scale requires highly efficient methods. Here, we developed a structural-alignment-based clustering algorithm—Foldseek cluster—that can cluster hundreds of m

Molecular BiologyBiochemistry, Genetics and Molecular Biology
13
Article|307 citations·2016
MMseqs software suite for fast and deep clustering and searching of large protein sequence sets
Maria Hauser, Martin Steinegger, Johannes Söding
SJR Q1BioinformaticsOA

MOTIVATION: Sequence databases are growing fast, challenging existing analysis pipelines. Reducing the redundancy of sequence databases by similarity clustering improves speed and sensitivity of iterative searches. But existing tools cannot efficiently cluster databases of the size of UniProt to 50% maximum pairwise sequence identity or below. Furthermore, in metagenomics experiments typically large fractions of reads cannot be matched to any known sequence anymore because searching with sensiti

Molecular BiologyBiochemistry, Genetics and Molecular Biology
14
Article|255 citations·2020
Terminating contamination: large-scale search identifies more than 2,000,000 contaminated entries in GenBank
Martin Steinegger, Steven L. Salzberg
SJR Q1Genome biologyOA

Genomic analyses are sensitive to contamination in public databases caused by incorrectly labeled reference sequences. Here, we describe Conterminator, an efficient method to detect and remove incorrectly labeled sequences by an exhaustive all-against-all sequence comparison. Our analysis reports contamination of 2,161,746, 114,035, and 14,148 sequences in the RefSeq, GenBank, and NR databases, respectively, spanning the whole range from draft to "complete" model organism genomes. Our method sca

Molecular BiologyBiochemistry, Genetics and Molecular Biology
15
Review|134 citations·2024
Easy and accurate protein structure prediction using ColabFold
Gyuri Kim, Sewon Lee, Eli Levy Karin, Hyunbin Kim, Yoshitaka Moriwaki, Sergey Ovchinnikov, Martin Steinegger, Milot Mirdita
SJR Q1Nature Protocols
Molecular BiologyBiochemistry, Genetics and Molecular Biology

Research Areas

Molecular BiologyEcologyMaterials ChemistryEpidemiologyBiotechnologyPlant Science

Martin Stiegeの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。