[論文レビュー] Benchmarking Transcriptomics Foundation Models for Perturbation Analysis : one PCA still rules them all
論文は、転写オミクス基盤モデルを撹乱分析のために公的データセットでベンチマークし、scVIとPCAが一般に基盤モデルを上回ることを示し、Structural Integrityを遺伝子活性の構造保持の新しい評価指標として導入している。
Understanding the relationships among genes, compounds, and their interactions in living organisms remains limited due to technological constraints and the complexity of biological data. Deep learning has shown promise in exploring these relationships using various data types. However, transcriptomics, which provides detailed insights into cellular states, is still underused due to its high noise levels and limited data availability. Recent advancements in transcriptomics sequencing provide new opportunities to uncover valuable insights, especially with the rise of many new foundation models for transcriptomics, yet no benchmark has been made to robustly evaluate the effectiveness of these rising models for perturbation analysis. This article presents a novel biologically motivated evaluation framework and a hierarchy of perturbation analysis tasks for comparing the performance of pretrained foundation models to each other and to more classical techniques of learning from transcriptomics data. We compile diverse public datasets from different sequencing techniques and cell lines to assess models performance. Our approach identifies scVI and PCA to be far better suited models for understanding biological perturbations in comparison to existing foundation models, especially in their application in real-world scenarios.
研究の動機と目的
- 撹乱分析のための生物学的根拠に基づくベンチマークの動機付け。
- 撹乱タスクに対して事前学習済みの転写オミクス基盤モデルを古典的手法と比較。
- データセットと技術に跨って撹乱シグナルを最もよく捉えるモデルを特定。
- 遺伝子活性の構造を保持する新しい評価指標としてStructural Integrityを導入。
提案手法
- 3つのシーケンス技術と複数の細胞系を含む多様な公的撹乱データセットを整備。
- 階層的な評価フレームワークを、指標としてiLISIバッチ統合、潜在表現の分離性(線形 probing)、撹乱の一貫性、局所潜在構造(kNN)、ゼロショット既知関係の再現、再構成の解釈性を定義。
- Structural Integrityを提案。中心対数発現とフロベニウス距離を用いた正規化ベースの指標で、バッチ内の撹乱構造の保持を定量化。
- モデル出力に後処理(コントロールベースのセンタリング、TVN、または生データ埋め込み)を適用し、モデルタスクごとに最も良い手法を選択。
- PCAとscVIをベンチマークし、Geneformer、scGPT、CellPLM、UCEなどの基盤モデルと比較してタスクごとに評価。
- 再現性のために完全な結果とコード(Tx-Evaluation)を提供。

実験結果
リサーチクエスチョン
- RQ1転写オミクス基盤モデルは、バッチ補正や分類以外の撹乱分析タスクへ一般化できるか。
- RQ2多様なデータセットとシーケンスモダリティにわたって、どのモデルと後処理戦略が撹乱効果を最もよく捉えるか。
- RQ3遺伝子分布仮定は撹乱タスクのモデル性能にどのように影響するか。
- RQ4PCAやscVIのような単純なモデルが、撹乱中心のベンチマークで複雑な基盤モデルを凌駕するか。
- RQ5撹乱表現を評価するStructural Integrity指標の有用性はどの程度か。
主な発見
- 基盤モデルはPCAとscVIに比べて撹乱タスクへ一般化が難しく、バッチ効果の低減を除きあまり高性能にはならない。
- scVIはゼロショット転送またはスクラッチ学習のいずれでも強力な性能とスケーラブルな結果を示し、多くの基盤モデルを上回る。
- 遺伝子分布(ZINB、NB、Poisson)はscVIの性能に実質的な影響を与え、データセット依存的である(Replogle対L1000)。
- scVIは少量の訓練データでも堅牢に学習し、単一細胞撹乱文脈でより多くのデータが増えるとスケールする。
- GeneformerやscGPTのような基盤モデルはバッチ効果低減に優れる一方、生物学的に意味のある撹乱タスクでは苦戦する。
- Structural Integrityは潜在的な遺伝子活性空間における撹乱関係の保存度を示す新しい指標である。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。