[논문 리뷰] DPLM-2: A Multimodal Diffusion Protein Language Model
DPLM-2는 이산 확산 단백질 언어 모델을 확장하여 조회 없이 구조 토크나이저와 다중 모달 학습 목표를 사용해 단백질 시퀀스와 구조를 함께 모델링하고 생성하며 무조건적 공동 생성 및 다양한 조건 생성 작업을 가능하게 한다.
Proteins are essential macromolecules defined by their amino acid sequences, which determine their three-dimensional structures and, consequently, their functions in all living organisms. Therefore, generative protein modeling necessitates a multimodal approach to simultaneously model, understand, and generate both sequences and structures. However, existing methods typically use separate models for each modality, limiting their ability to capture the intricate relationships between sequence and structure. This results in suboptimal performance in tasks that requires joint understanding and generation of both modalities. In this paper, we introduce DPLM-2, a multimodal protein foundation model that extends discrete diffusion protein language model (DPLM) to accommodate both sequences and structures. To enable structural learning with the language model, 3D coordinates are converted to discrete tokens using a lookup-free quantization-based tokenizer. By training on both experimental and high-quality synthetic structures, DPLM-2 learns the joint distribution of sequence and structure, as well as their marginals and conditionals. We also implement an efficient warm-up strategy to exploit the connection between large-scale evolutionary data and structural inductive biases from pre-trained sequence-based protein language models. Empirical evaluation shows that DPLM-2 can simultaneously generate highly compatible amino acid sequences and their corresponding 3D structures eliminating the need for a two-stage generation approach. Moreover, DPLM-2 demonstrates competitive performance in various conditional generation tasks, including folding, inverse folding, and scaffolding with multimodal motif inputs, as well as providing structure-aware representations for predictive tasks.
연구 동기 및 목표
- 단백질 시퀀스와 구조의 통합 모델링 필요성을 동기 부여하고 해결한다.
- 시퀀스와 구조의 결합 분포를 학습하는 다중 모달 프로틴 기초 모델을 개발한다.
- 3D 좌표를 언어 모델 학습을 위한 이산 토큰으로 변환하기 위해 구조 토크나이저를 활용한다.
- 구조 학습을 향상시키기 위해 사전 학습된 시퀀스 기반 지식을 사용해 워밍업한다.
- 구조 인식 표현을 사용한 무조건적 공동 생성 및 여러 조건부 생성 작업을 시연한다.
제안 방법
- 시퀀스와 구조를 통합 프레임워크에서 다루도록 이산 확산 단백질 언어 모델(DPLM)을 확장한다.
- 3D 백본 좌표를 이산 구조 토큰으로 토큰화하기 위해 조회 없이 양자화기(LFQ)를 도입한다.
- 구조 토큰을 아미노산 시퀀스와 연결하고 잔류물 수준의 위치를 공유 인코딩으로 정렬한다.
- 시퀀스 확산에서 노출 편향을 완화하기 위해 모달리티별 노이즈 스케줄러와 자기 혼합(self-mixup) 학습 전략을 적용한다.
- 선행 사전 학습된 시퀀스 기반 DPLM에서 LoRA를 사용한 효율적인 워밍업을 구현하여 사전 학습된 파라미터를 보존하면서 진화적 지식을 전달한다.
실험 결과
연구 질문
- RQ1하나의 단일 다중 모달 확산 모델이 고충실도로 단백질 시퀀스와 구조를 함께 모델링하고 생성할 수 있는가?
- RQ2언어 모델 프레임워크 내에서 구조 정보를 어떻게 효과적으로 학습할 수 있는가?
- RQ3접힘, 역접힘, 모티프 스캐폴딩 작업에 대한 다중 모달 조건화의 이점은 무엇인가?
- RQ4시퀀스 데이터에 대한 사전 학습과 데이터 증강이 다중 모달 생성과 다양성을 향상시키는가?
주요 결과
- DPLM-2는 2단계 연쇄 없이 호환 가능한 단백질 시퀀스와 3D 구조를 동시에 생성한다.
- 실험 데이터와 AlphaFold 예측 구조로 학습된 이 모델은 시퀀스와 구조에 대한 결합 분포, 주변 분포, 조건부 분포를 학습한다.
- DPLM-2는 다중 모달 입력으로 접힘, 역접힘, 모티프 스캐폴딩 작업에서 경쟁력 있는 성능을 보여준다.
- DPLM-2의 구조 인식 표현이 생성 외의 예측 작업을 향상시킨다.
- 사전 학습된 시퀀스 기반 DPLM과 데이터 증강의 워밍업은 특히 더 긴 단백질에서 설계 가능성과 다양성을 크게 향상시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.