Skip to main content
QUICK REVIEW

[논문 리뷰] GENERator: A Long-Context Generative Genomic Foundation Model

Wei Wu, Qiuyi Li|ArXiv.org|2025. 02. 11.
Gene expression and cancer classification인용 수 6
한 줄 요약

GENERATOR는 1.2B-parameter 생성적 게놈 기초 모델로, 98k base-pair 컨텍스트를 가지며, 386B bp의 진핵생물 DNA에서 학습되어 게놈 벤치마크에서 최첨단 성능을 달성하고 중심 도그마에 맞춘 단백질 코딩 및 프로모터 설계가 가능하다.

ABSTRACT

The rapid advancement of DNA sequencing has produced vast genomic datasets, yet interpreting and engineering genomic function remain fundamental challenges. Recent large language models have opened new avenues for genomic analysis, but existing approaches are often limited by restricted training scope, constrained generative capability, or prohibitive computational cost. We introduce GENErator, a generative genomic foundation model for long-context DNA modeling, with a context length of 98k nucleotides, pre-trained on 386 billion nucleotides of eukaryotic DNA. Without task-specific fine-tuning, GENERator exhibits strong intrinsic capabilities: unsupervised embedding analyses reveal phylogenetically coherent structure, and sequence recovery benchmarks demonstrate generative accuracy comparable to or exceeding state-of-the-art models with substantially improved computational efficiency. In a zero-shot setting, GENERator achieves competitive variant effect prediction performance relative to alignment-based methods, while remaining fully alignment-free and broadly applicable across species. With task-specific fine-tuning, the model attains leading performance on established genomic benchmarks. We further demonstrate practical generative applications. GENERator can generate protein-coding DNA sequences that translate into structurally plausible proteins and, through a prompt-guided design framework, design cis-regulatory elements with targeted activity profiles, including synthetic super-enhancers validated by high-throughput UMI-STARR-seq assays. Together, these results establish GENERator as an efficient and biologically grounded framework for genomic interpretation and programmable sequence design. Code and supplementary resources are available at https://github.com/GenerTeam/GENERator.

연구 동기 및 목표

  • DNA 데이터에 특화된 롱 컨텍스트 생성형 기초 모델로 게놈 서열 모델링을 발전시키다.
  • 확립된 벤치마크와 새로 제안된 게놈 벤치마크에서 최첨단 성능을 시연하다.
  • 중심 도그마에 맞춰 알려진 단백질 계열로 번역되는 단백질 코딩 서열을 생성함으로써 정합성을 보이다.
  • 활성 표적화를 포함한 프롬프트 반응형 프로모터 설계를 포함하여 서열 설계 능력을 탐구하다.
  • 롱-range 게놈 이해를 극대화하는 학습 전략과 토크나이저 선택을 조사하다.

제안 방법

  • Llama에서 영감을 받은 26층, 숨김 크기 2,048의 트랜스포머 디코더 아키텍처를 사용한다.
  • RefSeq의 진핵생물 DNA 386B 뉴클레오타이드를 대상으로 차기 토큰 예측(NTP)을 위한 6-mer 토크나이저로 학습한다.
  • 유전자 서열 학습과 전체 서열 학습을 비교하고 다운스트림 작업에 더 효과적인 의미론적으로 풍부한 영역을 식별한다.
  • 롱 컨텍스트 데이터를 효율적으로 처리하는 기술(Flash Attention, Zero Redundancy Optimizer)을 활용하고 난수화된 토크나이제이션 시작점을 도입하여 로버스트함을 향상시킨다.
  • Genomic Benchmarks, NT 작업, 그리고 새로운 Gener 작업(유전자/분류학 분류 및 다음-K-mer 예측 포함)을 평가하고 중심 도그마 및 프로모터 설계 작업을 분석한다.
  • 26층, 히든 사이즈 2048, 어휘 수 4128, 컨텍스트 길이 16384 토큰에 해당하는 98,304 bp를 나타내는 상세 아키텍처 규격과 학습 설정(배치 크기 2M 토큰, 6 에폭, AdamW, 코사인 워밍업)을 제공한다.
Figure 1: Overview of the Gener ator . (A) The pre-training dataset of the Gener ator encompasses a diverse range of eukaryotic organisms and gene types, totaling 386B nucleotides. (B) The pre-training employs the next token prediction (NTP) task, utilizing a 6-mer tokenizer. (C) Model comparison re
Figure 1: Overview of the Gener ator . (A) The pre-training dataset of the Gener ator encompasses a diverse range of eukaryotic organisms and gene types, totaling 386B nucleotides. (B) The pre-training employs the next token prediction (NTP) task, utilizing a 6-mer tokenizer. (C) Model comparison re

실험 결과

연구 질문

  • RQ1다양한 게놈 벤치마크와 작업에서 GENERator가 최첨단 성능을 달성할 수 있는가?
  • RQ2토크나이저 선택(6-mer)이 DNA 언어 모델에서의 다음 토큰 예측에 BPE나 단일 뉴클레오타이드 토크나이저와 비교해 어떤 영향을 미치는가?
  • RQ3유전자 영역(의미적으로 풍부한 데이터)에서의 학습이 다운스트림 게놈 작업에서 전체 게놈 학습보다 더 나은가?
  • RQ4모델이 중심 도그마 정합성을 보이는 단백질 코딩 DNA 서열을 생성하고 이를 단백질로 번역 가능한가?
  • RQ5프로모터 설계와 같은 시퀀스 설계에 GENERator가 얼마나 기여할 수 있는가(특정 활성이 목표로 하는 설계)?

주요 결과

  • Genomic Benchmarks, NT 작업, 그리고 새로 제안된 Gener 작업에서 최첨단 성능을 달성한다.
  • 1.2B 매개변수와 98k bp 컨텍스트를 사용하는 모델로, NT-multi, Enformer, GROVER, HyenaDNA, Caduceus 등의 기준선을 능가한다.
  • 유전자 서열 학습(의미적으로 풍부한 영역에 초점)이 다수의 분류군에서 다운스트림 작업에 대해 전체 서열 학습보다 우수하다.
  • 중심 도그마 정합성을 보이며 알려진 계통과 구조적으로 동등한 단백질로 번역되는 단백질 코딩 DNA 서열을 생성하고, 그 접힘성(AlphaFold) 및 분포형 당황도(Progen2)를 평가한다.
  • DeepSTARR 프로모터 데이터 세트를 사용한 프롬프트 반응형 활성이 있는 설계를 통해 프로모터 설계 능력을 보여주고, 시퀀스 최적화를 제어할 수 있게 한다.
Figure 2: Evaluation of next K-mer prediction. (A) Accuracy of the next K-mer prediction task across various tokenizers and input token lengths. (B) Comparison of the Gener ator against baseline models on a dataset comprised exclusively mammalian DNA.
Figure 2: Evaluation of next K-mer prediction. (A) Accuracy of the next K-mer prediction task across various tokenizers and input token lengths. (B) Comparison of the Gener ator against baseline models on a dataset comprised exclusively mammalian DNA.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.