[论文解读] ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics
ST-Align 是首个用于空间转录组学的多模态基础模型,通过三重目标对齐策略,在多个空间尺度上实现病理图像与基因表达数据的对齐。通过整合专用编码器、基于注意力的融合网络(ABFN)以及斑点与微环境之间的对比学习,其在六个数据集上的空间聚类和基因预测任务中均实现了最先进(SOTA)的零样本和少样本性能。
Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-level, the absence of a wider slide perspective and spatial intrinsic relationships limits their ability to capture ST-specific insights effectively. Here, we introduce ST-Align, the first foundation model designed for ST that deeply aligns image-gene pairs by incorporating spatial context, effectively bridging pathological imaging with genomic features. We design a novel pretraining framework with a three-target alignment strategy for ST-Align, enabling (1) multi-scale alignment across image-gene pairs, capturing both spot- and niche-level contexts for a comprehensive perspective, and (2) cross-level alignment of multimodal insights, connecting localized cellular characteristics and broader tissue architecture. Additionally, ST-Align employs specialized encoders tailored to distinct ST contexts, followed by an Attention-Based Fusion Network (ABFN) for enhanced multimodal fusion, effectively merging domain-shared knowledge with ST-specific insights from both pathological and genomic data. We pre-trained ST-Align on 1.3 million spot-niche pairs and evaluated its performance through two downstream tasks across six datasets, demonstrating superior zero-shot and few-shot capabilities. ST-Align highlights the potential for reducing the cost of ST and providing valuable insights into the distinction of critical compositions within human tissue.
研究动机与目标
- 为解决现有模型在捕捉空间转录组学(ST)数据中空间上下文和内在关系方面的局限性。
- 弥合高分辨率组织病理学图像与斑点和微环境水平上的全转录组基因表达谱之间的差距。
- 通过多模态预训练,开发一个可在少量微调下泛化于多样化ST数据集的基础模型。
- 通过增强的多模态融合与空间感知能力,提升空间领域识别与基因表达预测的准确性。
- 通过支持零样本和少样本迁移学习能力,降低空间转录组学的成本与复杂性。
提出的方法
- 提出一种三重目标对齐策略:(1)斑点级别图像-基因对齐,(2)微环境级别图像-基因对齐,(3)多模态特征的跨层级融合。
- 采用专用编码器——视觉Transformer用于图像,改进的Transformer用于基因序列,以捕捉不同尺度下的领域特定特征。
- 引入基于注意力的融合网络(ABFN),动态融合视觉与遗传嵌入,整合共享知识与ST特异性知识。
- 利用斑点-微环境对比损失($\mathcal{L}_{NS}$)将单个斑点与其更广泛的时空微环境对齐,增强空间上下文建模能力。
- 在573张人类组织切片的130万对斑点-微环境对上预训练ST-Align,涵盖正常、疾病和癌变样本。
- 采用多阶段预训练框架,结合对比学习与掩码自编码,以提升特征表示能力与鲁棒性。
实验结果
研究问题
- RQ1基础模型是否能有效实现空间转录组学中多空间尺度下的图像与基因模态对齐?
- RQ2在图像-基因对齐模型中,引入微环境级别空间上下文在多大程度上提升了性能?
- RQ3与简单拼接相比,基于注意力的融合网络(ABFN)在多模态特征融合方面提升了多少?
- RQ4ST-Align在空间聚类与基因预测等下游任务的零样本与少样本设置下的表现如何?
- RQ5斑点-微环境对比学习目标在ST数据的空间关系建模中起到了何种贡献?
主要发现
- 与其它多模态模型相比,ST-Align在预测非层状基因方面提升了+23.74%,凸显其在非结构特异性基因预测中的有效性。
- 在零样本空间聚类中,ST-Align优于CLIP与PLIP,能够准确区分人脑组织切片中L1与L2层之间的细微结构差异。
- 消融实验表明,移除自动编码器(AEs)与ABFN后,两项下游任务的性能分别下降8.06%与6.61%,证明二者在特征融合中的关键作用。
- 引入斑点-微环境对比损失($\mathcal{L}_{NS}$)使空间聚类性能提升17.76%,表明其在建模空间层级结构方面的价值。
- 与基线多模态模型相比,ST-Align在非层状基因预测任务中实现了+6.97%的性能提升,表明在基因特征上联合预训练的优势。
- 可视化结果证实,在零样本设置下,ST-Align比CLIP与PLIP更准确地勾勒出白质与L6层之间的边界。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。