[论文解读] Deep Learning in Spatially Resolved Transcriptomics: A Comprehensive Technical View
本文对空间分辨转录组学(SRT)中的深度学习方法进行了全面的技术综述,系统性地将21种前沿模型归类为六项分析任务——空间聚类、空间可变基因识别、细胞类型去卷积、基因表达填补、细胞间相互作用检测和空间区域映射。文章评估了其模型架构、优势、局限性及在SRT数据集上的性能表现,提供了一个统一的基准,并指出了在归一化、批次效应校正和多模态整合方面仍存在的开放性挑战。
Spatially resolved transcriptomics (SRT) has evolved rapidly through various technologies, enabling scientists to investigate both morphological contexts and gene expression profiling at single-cell resolution in parallel. SRT data are complex and multi-modal, comprising gene expression matrices, spatial information, and often high-resolution histology images. Because of this complexity and multi-modality, sophisticated computational algorithms are required to accurately analyze SRT data. Most efforts in this domain have been made to utilize conventional machine learning and statistical approaches, exhibiting sub-optimal results due to the complicated nature of SRT datasets. To address these shortcomings, researchers have recently employed deep learning algorithms including various state-of-the-art methods mainly in spatial clustering, spatially variable gene identification, and alignment. While great progress has been made in developing deep learning-based models for SRT data analysis, further improvement is still needed to create more biologically aware models that consider aspects such as phylogeny-aware clustering or the analysis of small histology image patches. Additionally, strategies for batch effect removal, normalization, and handling overdispersion and zero inflation patterns of gene expression are still needed in the analysis of SRT data using deep learning methods. In this paper, we provide a comprehensive overview of these deep learning methods, including their strengths and limitations. We also highlight new frontiers, current challenges, limitations, and open questions in this field. Also, we provide a comprehensive list of all available SRT databases that can be used as an extensive resource for future studies.
研究动机与目标
- 提供空间分辨转录组学(SRT)数据中应用的深度学习方法的系统性技术概述。
- 根据其分析任务(如空间聚类、基因表达填补和细胞类型去卷积)对21种深度学习模型进行分类与比较。
- 评估深度学习模型相较于传统机器学习和统计方法在SRT分析中的优势与局限性。
- 识别SRT数据分析中的开放性挑战,包括批次效应校正、归一化、过度分散和零膨胀处理。
- 整理一份公开可获取的SRT数据库完整列表,以支持未来研究与模型基准测试。
提出的方法
- 对SRT相关研究中发表的21种深度学习模型进行系统性文献综述与分类。
- 将模型归类为六项主要分析任务:空间聚类、空间可变基因(SVG)识别、细胞类型去卷积、基因表达填补、细胞间相互作用检测和空间区域映射。
- 分析模型架构,包括卷积神经网络(CNNs)、图卷积网络(GCNs)、变分自编码器(VAEs)和深度神经网络(DNNs),重点分析其对空间、组织学和基因表达数据的利用方式。
- 使用标准化指标(如ARI、F1-score、AUC和重建误差)评估模型性能,若已报告则纳入评估。
- 在补充材料中整合关键模型的数学公式(如CNNTL中的三元组损失、conST中的对比学习)。
- 通过注意力机制、自编码器和基于图的学习方法,实现多模态数据(基因表达、组织学图像、空间坐标)的整合。
实验结果
研究问题
- RQ1深度学习模型在分析复杂、多模态SRT数据时,相较于传统机器学习和统计方法有何改进?
- RQ2用于空间转录组学任务的深度学习模型在架构与方法论上存在哪些关键差异?
- RQ3当前的深度学习模型在多大程度上考虑了生物先验知识,如系统发育关系、组织形态或空间邻近关系?
- RQ4现有深度学习方法在SRT分析中的主要局限性是什么,特别是在数据归一化、批次效应校正以及零膨胀基因表达处理方面?
- RQ5如何在SRT的深度学习框架中优化组织学图像、空间坐标与基因表达数据的多模态整合?
主要发现
- 在空间聚类、SVG检测和细胞类型去卷积等任务中,深度学习模型显著优于传统机器学习和统计方法,尤其在利用空间上下文和组织学图像时表现更优。
- 如conST和XFuse等模型通过对比学习和变分自编码器分别在空间区域检测和基因表达填补任务中达到最先进性能。
- 尽管已有进展,许多模型仍受限于对预训练ImageNet权重的依赖或缺乏领域特定的预训练,导致其在SRT数据上的泛化能力受限。
- Tangram和JSTA等模型的性能高度依赖于标注的解剖模板(如CCF)的可用性,限制了其在非小鼠脑组织中的适用性。
- 在处理基因表达数据中的技术性偏差(如过度分散和零膨胀)方面仍存在显著缺口,仅有少数深度学习模型在其损失函数中显式整合了这些特征。
- 本文整理了21个SRT数据库和数据集的完整列表,包括10x Visium、MERFISH、seqFISH、Stereo-seq和Slide-seq,为模型训练与基准测试提供了关键资源。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。