[논문 리뷰] LangCell: Language-Cell Pre-training for Cell Identity Understanding
LangCell는 단일세포 전사체 데이터와 세포 정체성에 대한 자연어 기술을 통합하는 최초의 언어-세포 사전학습 프레임워크이다. 다중 작업 학습 방식을 사용하여 마스킹 유전자 모델링, 세포-세포 대비 학습, 세포-텍스트 대비 학습 및 매칭을 동시에 사전학습한다. 이는 세포 정체성 이해 분야에서 zero-shot, few-shot, fine-tuning 환경 모두에서 최고 성능을 기록하며, 다양한 벤치마크에서 기존 모델들을 능가한다.
Cell identity encompasses various semantic aspects of a cell, including cell type, pathway information, disease information, and more, which are essential for biologists to gain insights into its biological characteristics. Understanding cell identity from the transcriptomic data, such as annotating cell types, has become an important task in bioinformatics. As these semantic aspects are determined by human experts, it is impossible for AI models to effectively carry out cell identity understanding tasks without the supervision signals provided by single-cell and label pairs. The single-cell pre-trained language models (PLMs) currently used for this task are trained only on a single modality, transcriptomics data, lack an understanding of cell identity knowledge. As a result, they have to be fine-tuned for downstream tasks and struggle when lacking labeled data with the desired semantic labels. To address this issue, we propose an innovative solution by constructing a unified representation of single-cell data and natural language during the pre-training phase, allowing the model to directly incorporate insights related to cell identity. More specifically, we introduce $ extbf{LangCell}$, the first $ extbf{Lang}$uage-$ extbf{Cell}$ pre-training framework. LangCell utilizes texts enriched with cell identity information to gain a profound comprehension of cross-modal knowledge. Results from experiments conducted on different benchmarks show that LangCell is the only single-cell PLM that can work effectively in zero-shot cell identity understanding scenarios, and also significantly outperforms existing models in few-shot and fine-tuning cell identity understanding scenarios.
연구 동기 및 목표
- 라벨이 없는 데이터에서 단일 모odal 사전학습 언어 모델이 세포 정체성을 이해하는 데 한계가 있음을 해결하기 위해.
- 세포 유형, 질병, 경로 정보와 같은 인간이 주석 처리한 세포 정체성의 의미적 지식을 단일세포 데이터의 표현 학습에 통합하기 위해.
- 직접적인 zero-shot 추론을 가능하게 하는 통합형 다중모달 프레임워크를 개발하기 위해.
- 세포 유형 주석 및 배치 통합과 같은 후행 작업에서 비용이 많이 들고 부족한 라벨 데이터에 대한 의존도를 줄이기 위해.
- 공동 사전학습을 지원하기 위해 2750만 건의 항목을 포함한 대규모 고품질의 세포-텍스트 데이터셋(scLibrary)을 구축하기 위해.
제안 방법
- OBO Foundry에서 확보한 세포 정체성 측면에 대한 자연어 기술과 연계된 2750만 건의 단일세포 RNA-seq 항목을 포함한 대규모 데이터셋인 scLibrary를 구축하였다.
- 마스킹 유전자 모델링(MGM), 세포-세포 대비 학습(C-C), 세포-텍스트 대비 학습(C-T), 세포-텍스트 매칭(CTM)의 네 가지 작업을 통합한 다중 작업 사전학습 목표를 설계하였다.
- 변환기 기반 아키텍처를 사용하여 단일세포 RNA-seq 데이터와 텍스트 기술을 동일한 표현 공간에 함께 인코딩하였다.
- 마스킹 유전자 모델링을 통해 유전자 공표현 패턴을 학습하고 전사체 표현을 향상시켰다.
- 다양한 모달 간 표현을 정렬하기 위해 세포 간 대비 학습(C-C)과 세포-텍스트 간 대비 학습(C-T)을 적용하였다.
- 제한된 라벨 데이터를 사용하여 후행 작업에 모델을 미세조정하였으며, 동시에 라벨 없는 상태에서의 zero-shot 성능도 평가하였다.

실험 결과
연구 질문
- RQ1사전학습 모델이 단일세포 전사체 데이터와 자연어 기술을 동시에 학습하여 zero-shot 세포 정체성 이해를 달성할 수 있는가?
- RQ2단일 모달 모델 대비 언어-세포 공동 사전학습이 세포 유형 주석 및 배치 통합 작업에서 성능을 얼마나 향상시키는가?
- RQ3라벨이 없는 미세조정 데이터 없이도 모델이 새로운 세포 유형이나 질병에 얼마나 잘 일반화되는가?
- RQ4온톨로지에서 유래한 인간이 수작업으로 확보한 의미적 지식(예: 온톨로지 기반)을 사전학습 과정에 통합할 경우 어떤 영향을 미치는가?
- RQ5다중 모달 대비 목표를 포함한 다중 작업 학습이 세포 정체성 작업의 표현 학습을 어떻게 향상시키는가?
주요 결과
- LangCell는 PBMC10K, PBMC3&68K, Zheng68K, macParland, Aizarani, Perirhinal Cortex, Tabula Sapiens 등 다양한 벤치마크에서 zero-shot, few-shot, 전체 미세조정 설정에서 최고 성능을 기록하였다.
- LangCell는 효과적인 zero-shot 세포 유형 주석이 가능한 최초의 단일세포 PLM이며, 대부분의 경우 few-shot 기준선을 능가하였다.
- Tabula Sapiens 데이터셋에서 LangCell는 매크로파지 및 B 세포를 포함한 상위 10개 세포 유형에서 높은 재현율을 기록하며 강력한 세포-텍스트 검색 성능를 보였다.
- 특히 라벨 데이터가 부족한 저자료 환경에서 기존 모델들보다 뚜렷이 뛰어난 성능을 보이며 세포 유형 주석 작업에서 뛰어난 성능을 보였다.
- 제거 실험을 통해 MGM, C-C, C-T, CTM의 네 가지 사전학습 목표 모두가 성능 향상에 기여했으며, 특히 C-T와 CTM가 다중 모달 정렬에 효과적이었다.
- LangCell는 새로운 세포 유형과 질병으로의 일반화 능력이 뛰어나 의미적 지식 통합으로 인한 강력한 전이 학습 능력을 보였다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.