Skip to main content
QUICK REVIEW

[논문 리뷰] Benchmarking Transcriptomics Foundation Models for Perturbation Analysis : one PCA still rules them all

Ihab Bendidi, Shawn Whitfield|arXiv (Cornell University)|2024. 10. 17.
RNA Research and Splicing인용 수 8
한 줄 요약

본 논문은 perturbation 분석을 위한 전사체 기초 모델의 벤치마크를 공개 데이터셋에서 수행하고, 일반적으로 scVI와 PCA가 foundation-model보다 우수하다는 것을 발견하며, Structural Integrity를 새로운 평가 지표로 제안합니다.

ABSTRACT

Understanding the relationships among genes, compounds, and their interactions in living organisms remains limited due to technological constraints and the complexity of biological data. Deep learning has shown promise in exploring these relationships using various data types. However, transcriptomics, which provides detailed insights into cellular states, is still underused due to its high noise levels and limited data availability. Recent advancements in transcriptomics sequencing provide new opportunities to uncover valuable insights, especially with the rise of many new foundation models for transcriptomics, yet no benchmark has been made to robustly evaluate the effectiveness of these rising models for perturbation analysis. This article presents a novel biologically motivated evaluation framework and a hierarchy of perturbation analysis tasks for comparing the performance of pretrained foundation models to each other and to more classical techniques of learning from transcriptomics data. We compile diverse public datasets from different sequencing techniques and cell lines to assess models performance. Our approach identifies scVI and PCA to be far better suited models for understanding biological perturbations in comparison to existing foundation models, especially in their application in real-world scenarios.

연구 동기 및 목표

  • 전사체학에서 perturbation 분석을 위한 생물학적으로 근거 있는 벤치마크를 제안한다.
  • perturbation 작업에서 사전 학습된 전사체 기초 모델을 고전적 방법과 비교한다.
  • 데이터세트와 기법 전반에서 어떤 모델이 perturbation 신호를 가장 잘 포착하는지 식별한다.
  • 유전자 활성 구조 보존을 위한 새로운 평가 지표로 Structural Integrity를 도입한다.

제안 방법

  • 세 가지 시퀀싱 기법과 여러 세포주에서 다양한 공개 perturbation 데이터셋을 선별한다.
  • iLISI 배치 통합, 잠재 분리도(linear probing), perturbation 일관성, 국소 잠재 구조(kNN), zero-shot 알려진 관계 재현, 재구성 해석 가능성 등의 지표를 포함하는 계층적 평가 프레임워크를 정의한다.
  • Centered log-expression과 Frobenius 거리를 이용한 정규화 기반 지표인 Structural Integrity를 제안하여 배치 내 perturbation 구조의 보존 정도를 정량화한다.
  • 모델 출력에 대해 컨트롤 기반 센터링, TVN, 또는 원시 임베딩과 같은 후처리를 적용하고 모델-작업별로 최적의 방법을 선택한다.
  • baseline PCA와 scVI를 Geneformer, scGPT, CellPLM, UCE 등과 같은 foundation 모델과 각 작업에서 벤치마킹한다.
  • 재현성을 위해 전체 결과와 코드(Tx-Evaluation)를 제공한다.
Figure 1: Known biological relationship recall scores for (Replogle et al., 2022 ) and L1000 Assay for scVI, trained using different gene distributions. Different datasets and sequencing approaches benefit from different gene distributions.
Figure 1: Known biological relationship recall scores for (Replogle et al., 2022 ) and L1000 Assay for scVI, trained using different gene distributions. Different datasets and sequencing approaches benefit from different gene distributions.

실험 결과

연구 질문

  • RQ1전사체학 기초 모델이 배치 보정이나 분류를 넘어 perturbation 분석 작업에 일반화될 수 있는가?
  • RQ2다양한 데이터 세트와 시퀀싱 모듈성에 걸쳐 어떤 모델과 후처리 전략이 perturbation 효과를 가장 잘 포착하는가?
  • RQ3유전자 분포 가정이 perturbation 작업에서 모델 성능에 어떤 영향을 미치는가?
  • RQ4PCA나 scVI 같은 단순한 모델이 perturbation 중심 벤치마크에서 복잡한 기초 모델보다 우수한가?
  • RQ5Perturbation 표현 평가를 위한 Structural Integrity 지표의 유용성은 무엇인가?

주요 결과

  • 기초 모델은 배치 효과 감소를 제외하고 perturbation 작업에 대해 PCA와 scVI에 비해 일반화가 좋지 않다.
  • scVI(처음부터 학습하거나 제로샷 전이로 학습)는 종종 강한 성능과 확장 가능한 결과를 달성하여 많은 기초 모델보다 우수하다.
  • 유전자 분포(ZINB, NB, Poisson)가 scVI 성능에 실질적으로 영향을 주며 데이터세트에 의존적이다(Replogle 대 L1000).
  • scVI는 적은 학습 데이터에서도 견고한 학습을 보이며 단일세포 perturbation 맥락에서 더 많은 데이터와 함께 확장된다.
  • Geneformer 및 scGPT와 같은 기초 모델은 주로 배치 효과 감소에 뛰어나고 생물학적으로 의미 있는 perturbation 작업에서는 어려움을 보인다.
  • Structural Integrity는 잠재 유전자 활성 공간에서 perturbation 관계가 얼마나 잘 보존되는지 나타내는 새로운 지표이다.
Figure 2: Training scVI on different amounts of samples, using distinct batches and perturbations, before evaluating known biological relationship retrieval. scVI is robust to very low data regime, and shows strong data scaling laws on single cell data.
Figure 2: Training scVI on different amounts of samples, using distinct batches and perturbations, before evaluating known biological relationship retrieval. scVI is robust to very low data regime, and shows strong data scaling laws on single cell data.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.