Skip to main content

마틴 슈타이네거 교수

Martin Stiege

서울대학교 · 생화학·유전·분자생물학

연구실 소개

마틴 슈타이네거 교수의 연구실은 단백질 구조 예측과 시퀀스 분석을 위한 고속 알고리즘 및 소프트웨어 개발에 주력하고 있습니다. 특히 MMseqs2 기반의 빠른 서열 검색과 ColabFold를 통해 GPU 기반으로 하루에 수천 개의 단백질 구조를 예측할 수 있는 혁신적인 플랫폼을 제공합니다. 환경 샘플에서 유래한 단백질 데이터베이스 구축을 통해 메타게놈 분석과 기능 예측의 가능성을 넓히고 있으며, 대량의 단백질 서열을 효율적으로 클러스터링하는 Linclust 알고리즘 개발로도 주목받고 있습니다. 이는 생물정보학 및 구조생물학 분야의 대규모 데이터 분석을 혁신적으로 가속화합니다.

단백질 구조 예측고속 서열 검색메타게놈 분석대량 클러스터링생물정보학 소프트웨어

연구 현황

논문 수
175
총 인용 수
74,791
최근 5년 논문
111
주요 분야
생화학·유전·분자생물학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
111총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
16,323총합
20222023202420252026

주요 논문

15
1
논문|인용수 9,275·2022
ColabFold: making protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, Martin Steinegger
SJR Q1FWCI 772.4Nature MethodsOA

ColabFold offers accelerated prediction of protein structures and complexes by combining the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. ColabFold's 40-60-fold faster search and optimized model utilization enables prediction of close to 1,000 structures per day on a server with one graphics processing unit. Coupled with Google Colaboratory, ColabFold becomes a free and accessible platform for protein folding. ColabFold is open-source software available at https://github.com/s

Molecular BiologyBiochemistry, Genetics and Molecular Biology
2
논문|인용수 4,806·2017
MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets
Martin Steinegger, Johannes Söding
SJR Q1FWCI 33.7Nature Biotechnology
Molecular BiologyBiochemistry, Genetics and Molecular Biology
3
논문|인용수 1,244·2019
HH-suite3 for fast remote homology detection and deep protein annotation
Martin Steinegger, Markus Meier, Milot Mirdita, Harald Vöhringer, Stephan J. Haunsberger, Johannes Söding
SJR Q1FWCI 43.6BMC BioinformaticsOA

The added functionalities and increased speed of HHsearch and HHblits should facilitate their use in large-scale protein structure and function prediction, e.g. in metagenomics and genomics projects.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
4
논문|인용수 958·2018
Clustering huge protein sequence sets in linear time
Martin Steinegger, Johannes Söding
SJR Q1FWCI 24.5Nature CommunicationsOA

Metagenomic datasets contain billions of protein sequences that could greatly enhance large-scale functional annotation and structure prediction. Utilizing this enormous resource would require reducing its redundancy by similarity clustering. However, clustering hundreds of millions of sequences is impractical using current algorithms because their runtimes scale as the input set size N times the number of clusters K, which is typically of similar order as N, resulting in runtimes that increase

Molecular BiologyBiochemistry, Genetics and Molecular Biology
5
논문|인용수 886·2016
Uniclust databases of clustered and deeply annotated protein sequences and alignments
Milot Mirdita, Lars von den Driesch, Clovis Galiez, María Martin, Johannes Söding, Martin Steinegger
SJR Q1FWCI 8.7Nucleic Acids ResearchOA

We present three clustered protein sequence databases, Uniclust90, Uniclust50, Uniclust30 and three databases of multiple sequence alignments (MSAs), Uniboost10, Uniboost20 and Uniboost30, as a resource for protein sequence analysis, function prediction and sequence searches. The Uniclust databases cluster UniProtKB sequences at the level of 90%, 50% and 30% pairwise sequence identity. Uniclust90 and Uniclust50 clusters showed better consistency of functional annotation than those of UniRef90 an

Molecular BiologyBiochemistry, Genetics and Molecular Biology
6
리뷰|인용수 668·2022
Metagenome analysis using the Kraken software suite
Jennifer Lu, Natalia Rincon, Derrick E. Wood, Florian P. Breitwieser, Christopher Pockrandt, Ben Langmead, Steven L. Salzberg, Martin Steinegger
SJR Q1FWCI 54.8Nature ProtocolsOA
Molecular BiologyBiochemistry, Genetics and Molecular Biology
7
논문|인용수 660·2019
MMseqs2 desktop and local web server app for fast, interactive sequence searches
Milot Mirdita, Martin Steinegger, Johannes Söding
SJR Q1FWCI 14.6BioinformaticsOA

Supplementary data are available at Bioinformatics online.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
8
preprint|인용수 586·2021
ColabFold - Making protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, Martin Steinegger
bioRxiv (Cold Spring Harbor Laboratory)OA

ColabFold offers accelerated protein structure and complex predictions by combining the fast homology search of MMseqs2 with AlphaFold2 or RoseTTAFold. ColabFold’s 40 - 60× faster search and optimized model use allows predicting close to a thousand structures per day on a server with one GPU. Coupled with Google Colaboratory, ColabFold becomes a free and accessible platform for protein folding. ColabFold is open-source software available at github.com/sokrypton/ColabFold . Its novel environmenta

Molecular BiologyBiochemistry, Genetics and Molecular Biology
9
논문|인용수 471·2019
Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold
Martin Steinegger, Milot Mirdita, Johannes Söding
SJR Q1FWCI 15.0Nature MethodsOA

The open-source de novo protein-level assembler, Plass ( https://plass.mmseqs.com ), assembles six-frame-translated sequencing reads into protein sequences. It recovers 2-10 times more protein sequences from complex metagenomes and can assemble huge datasets. We assembled two redundancy-filtered reference protein catalogs, 2 billion sequences from 640 soil samples (soil reference protein catalog) and 292 million sequences from 775 marine eukaryotic metatranscriptomes (marine eukaryotic reference

Molecular BiologyBiochemistry, Genetics and Molecular Biology
10
preprint|인용수 377·2022
Fast and accurate protein structure search with Foldseek
Michel van Kempen, Stephanie Kim, Charlotte Tumescheit, Milot Mirdita, Jeong-Jae Lee, Cameron L. M. Gilchrist, Johannes Söding, Martin Steinegger
bioRxiv (Cold Spring Harbor Laboratory)OA

As structure prediction methods are generating millions of publicly available protein structures, searching these databases is becoming a bottleneck. Foldseek aligns the structure of a query protein against a database by describing the amino acid backbone of proteins as sequences over a structural alphabet. Foldseek decreases computation times by four to five orders of magnitude with 86%, 88% and 133% of the sensitivities of DALI, TM-align and CE, respectively.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
11
논문|인용수 337·2023
Clustering predicted structures at the scale of the known protein universe
Inigo Barrio‐Hernandez, Jingi Yeo, Jürgen Jänes, Milot Mirdita, Cameron L. M. Gilchrist, Tanita Wein, Mihály Váradi, Sameer Velankar, Pedro Beltrão, Martin Steinegger
SJR Q1FWCI 51.7NatureOA

Proteins are key to all cellular processes and their structure is important in understanding their function and evolution. Sequence-based predictions of protein structures have increased in accuracy<sup>1</sup>, and over 214 million predicted structures are available in the AlphaFold database<sup>2</sup>. However, studying protein structures at this scale requires highly efficient methods. Here, we developed a structural-alignment-based clustering algorithm-Foldseek cluster-that can cluster hund

Molecular BiologyBiochemistry, Genetics and Molecular Biology
12
논문|인용수 294·2016
MMseqs software suite for fast and deep clustering and searching of large protein sequence sets
Maria Hauser, Martin Steinegger, Johannes Söding
SJR Q1FWCI 3.5BioinformaticsOA

Supplementary data are available at Bioinformatics online.

Molecular BiologyBiochemistry, Genetics and Molecular Biology
13
논문|인용수 251·2020
Terminating contamination: large-scale search identifies more than 2,000,000 contaminated entries in GenBank
Martin Steinegger, Steven L. Salzberg
SJR Q1FWCI 13.3Genome biologyOA

Abstract Genomic analyses are sensitive to contamination in public databases caused by incorrectly labeled reference sequences. Here, we describe Conterminator, an efficient method to detect and remove incorrectly labeled sequences by an exhaustive all-against-all sequence comparison. Our analysis reports contamination of 2,161,746, 114,035, and 14,148 sequences in the RefSeq, GenBank, and NR databases, respectively, spanning the whole range from draft to “complete” model organism genomes. Our m

Molecular BiologyBiochemistry, Genetics and Molecular Biology
14
preprint|인용수 88·2019
HH-suite3 for fast remote homology detection and deep protein annotation
Martin Steinegger, Markus Meier, Milot Mirdita, Harald Vöhringer, Stephan J. Haunsberger, Johannes Söding
bioRxiv (Cold Spring Harbor Laboratory)OA

Abstract Background HH-suite is a widely used open source software suite for sensitive sequence similarity searches and protein fold recognition. It is based on pairwise alignment of profile Hidden Markov models (HMMs), which represent multiple sequence alignments of homologous sequences. Results We developed a single-instruction multiple-data (SIMD) vectorized implementation of the Viterbi algorithm for profile HMM alignment and introduced various other speed-ups. This accelerated HHsearch by a

Molecular BiologyBiochemistry, Genetics and Molecular Biology
15
preprint|인용수 88·2024
Multiple Protein Structure Alignment at Scale with FoldMason
Cameron L. M. Gilchrist, Milot Mirdita, Martin Steinegger
bioRxiv (Cold Spring Harbor Laboratory)OA

Abstract Protein structure is conserved beyond sequence, making multiple structural alignment (MSTA) essential for analyzing distantly related proteins. Computational prediction methods have vastly extended our repository of available proteins structures, requiring fast and accurate MSTA methods. Here, we introduce FoldMason, a progressive MSTA method that leverages the structural alphabet from Foldseek, a pairwise structural aligner, for multiple alignment of hundreds of thousands of protein st

Molecular BiologyBiochemistry, Genetics and Molecular Biology

대표 연구 분야

Molecular BiologyEcologyMaterials ChemistryEpidemiologyPlant ScienceFood Science

마틴 슈타이네거 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.