Skip to main content
QUICK REVIEW

[Paper Review] ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics

Yuxiang Lin, Ling Luo|arXiv (Cornell University)|Nov 25, 2024
Single-cell and spatial transcriptomicsBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

ST-Align is the first multimodal foundation model for spatial transcriptomics that aligns pathological images with gene expression data across multiple spatial scales using a three-target alignment strategy. By integrating specialized encoders, an Attention-Based Fusion Network (ABFN), and contrastive learning between spots and niches, it achieves state-of-the-art zero-shot and few-shot performance in spatial clustering and gene prediction across six datasets.

ABSTRACT

Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-level, the absence of a wider slide perspective and spatial intrinsic relationships limits their ability to capture ST-specific insights effectively. Here, we introduce ST-Align, the first foundation model designed for ST that deeply aligns image-gene pairs by incorporating spatial context, effectively bridging pathological imaging with genomic features. We design a novel pretraining framework with a three-target alignment strategy for ST-Align, enabling (1) multi-scale alignment across image-gene pairs, capturing both spot- and niche-level contexts for a comprehensive perspective, and (2) cross-level alignment of multimodal insights, connecting localized cellular characteristics and broader tissue architecture. Additionally, ST-Align employs specialized encoders tailored to distinct ST contexts, followed by an Attention-Based Fusion Network (ABFN) for enhanced multimodal fusion, effectively merging domain-shared knowledge with ST-specific insights from both pathological and genomic data. We pre-trained ST-Align on 1.3 million spot-niche pairs and evaluated its performance through two downstream tasks across six datasets, demonstrating superior zero-shot and few-shot capabilities. ST-Align highlights the potential for reducing the cost of ST and providing valuable insights into the distinction of critical compositions within human tissue.

Motivation & Objective

  • To address the limitation of existing models in capturing spatial context and intrinsic relationships in spatial transcriptomics (ST) data.
  • To bridge the gap between high-resolution histopathological images and whole-transcriptomic gene expression profiles at the spot and niche levels.
  • To develop a foundation model that generalizes across diverse ST datasets with minimal fine-tuning by leveraging multimodal pretraining.
  • To improve the accuracy of spatial domain identification and gene expression prediction through enhanced multimodal fusion and spatial perception.
  • To reduce the cost and complexity of spatial transcriptomics by enabling zero-shot and few-shot transfer learning capabilities.

Proposed method

  • Proposes a three-target alignment strategy: (1) spot-level image-gene alignment, (2) niche-level image-gene alignment, and (3) cross-level fusion of multimodal features.
  • Employs specialized encoders—Vision Transformer for images and a modified Transformer for gene sequences—to capture domain-specific features at different scales.
  • Introduces an Attention-Based Fusion Network (ABFN) to dynamically fuse visual and genetic embeddings, integrating both shared and ST-specific knowledge.
  • Utilizes a spot-niche contrastive loss ($\mathcal{L}_{NS}$) to align individual spots with their broader spatial niches, enhancing spatial context modeling.
  • Pretrains ST-Align on 1.3 million spot-niche pairs from 573 human tissue slices, including normal, diseased, and cancerous samples.
  • Applies a multi-stage pretraining framework combining contrastive learning and masked autoencoding to improve feature representation and robustness.

Experimental results

Research questions

  • RQ1Can a foundation model effectively align image and gene modalities across multiple spatial scales in spatial transcriptomics?
  • RQ2How does incorporating niche-level spatial context improve the performance of image-gene alignment models?
  • RQ3To what extent does the Attention-Based Fusion Network (ABFN) enhance multimodal feature integration compared to simple concatenation?
  • RQ4How does ST-Align perform in zero-shot and few-shot settings for downstream tasks like spatial clustering and gene prediction?
  • RQ5What is the contribution of the spot-niche contrastive learning objective to spatial relationship modeling in ST data?

Key findings

  • ST-Align achieved a +23.74% improvement in predicting Non-Laminar Genes compared to other multimodal models, highlighting its effectiveness in non-structure-specific gene prediction.
  • In zero-shot spatial clustering, ST-Align outperformed CLIP and PLIP by accurately distinguishing subtle structural differences between layers L1 and L2 in human brain tissue slices.
  • The ablation study showed that removing the AEs and ABFN reduced performance by 8.06% and 6.61% in the two downstream tasks, proving their critical role in feature fusion.
  • Incorporating the spot-niche contrastive loss ($\mathcal{L}_{NS}$) improved performance by 17.76% in spatial clustering, demonstrating its value in modeling spatial hierarchies.
  • ST-Align achieved a +6.97% improvement in Non-Laminar Gene prediction over baseline multimodal models, indicating the benefit of joint pretraining on genetic features.
  • Visualization results confirmed that ST-Align better delineates boundaries between white matter and layer L6 compared to CLIP and PLIP in a zero-shot setting.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.