Skip to main content
QUICK REVIEW

[论文解读] Challenges and opportunities to computationally deconvolve heterogeneous tissue with varying cell sizes using single cell RNA-sequencing datasets

Sean K. Maden, Sang Ho Kwon|arXiv (Cornell University)|May 10, 2023
Single-cell and spatial transcriptomics被引用 4
一句话总结

本文指出,现有用于整体RNA-seq数据的去卷积方法在应用于细胞大小高度可变的组织(如大脑或免疫组织)时会失效,因为这些方法将细胞大小与转录组活性混淆,导致细胞比例估计不准确。作者主张建立来自相同组织块的标准化、多组学‘金标准’数据集,以促进开发和验证稳健的、具备细胞大小感知能力的去卷积方法。

ABSTRACT

Deconvolution of cell mixtures in "bulk" transcriptomic samples from homogenate human tissue is important for understanding the pathologies of diseases. However, several experimental and computational challenges remain in developing and implementing transcriptomics-based deconvolution approaches, especially those using a single cell/nuclei RNA-seq reference atlas, which are becoming rapidly available across many tissues. Notably, deconvolution algorithms are frequently developed using samples from tissues with similar cell sizes. However, brain tissue or immune cell populations have cell types with substantially different cell sizes, total mRNA expression, and transcriptional activity. When existing deconvolution approaches are applied to these tissues, these systematic differences in cell sizes and transcriptomic activity confound accurate cell proportion estimates and instead may quantify total mRNA content. Furthermore, there is a lack of standard reference atlases and computational approaches to facilitate integrative analyses, including not only bulk and single cell/nuclei RNA-seq data, but also new data modalities from spatial -omic or imaging approaches. New multi-assay datasets need to be collected with orthogonal data types generated from the same tissue block and the same individual, to serve as a "gold standard" for evaluating new and existing deconvolution methods. Below, we discuss these key challenges and how they can be addressed with the acquisition of new datasets and approaches to analysis.

研究动机与目标

  • 识别并解决在整体组织样本转录组去卷积中细胞大小异质性的关键挑战。
  • 强调当前去卷积算法在训练时基于细胞大小相近的组织,当应用于大脑或免疫系统等细胞大小存在极端差异的组织时会产生偏差估计。
  • 强调缺乏标准化参考图谱和整合的多组学数据集,以验证去卷积方法。
  • 倡导从相同组织块生成正交的、多模态数据集(如整体RNA-seq、单核RNA-seq、空间组学)作为金标准。
  • 推动开发能够显式考虑细胞大小和总mRNA含量的新型计算方法,以提高细胞比例推断的准确性。

提出的方法

  • 提出从相同组织块构建多组学数据集,整合整体RNA-seq、单细胞/单核RNA-seq以及空间组学或成像数据。
  • 将这些整合数据集用作‘金标准’参考,以基准测试和验证现有及新型去卷积算法。
  • 开发显式将细胞大小和总mRNA含量建模为去卷积中混杂变量的计算方法。
  • 将现有去卷积算法应用于多种组织(如大脑、免疫组织)以展示因大小差异导致的系统性偏差。
  • 强调需要构建不仅包含转录组谱,还包含形态测量和亚细胞数据的参考图谱。
  • 鼓励使用正交数据模态以在去卷积中区分生物信号与技术伪影。

实验结果

研究问题

  • RQ1细胞大小和总mRNA含量的差异在多大程度上影响当前去卷积算法在整体RNA-seq分析中的准确性?
  • RQ2当应用于异质性组织时,现有去卷积方法在多大程度上会将细胞大小与细胞类型比例混淆?
  • RQ3当前单细胞/单核RNA-seq参考图谱在实现对细胞大小极端变异组织的准确去卷积方面存在哪些关键局限?
  • RQ4如何利用来自相同组织块的多组学数据集作为评估去卷积方法的金标准?
  • RQ5需要何种计算框架才能在去卷积中将细胞大小效应与真实的生物细胞类型组成分离开来?

主要发现

  • 现有去卷积方法在细胞大小高度可变的组织(如大脑和免疫组织)中系统性地错误估计细胞比例。
  • 偏差的产生是因为算法将由细胞大小引起的总mRNA含量差异与实际的细胞类型丰度混淆。
  • 当前的单细胞/单核RNA-seq参考图谱不足以在细胞大小极端异质的组织中实现准确的去卷积。
  • 缺乏标准化的、多模态参考数据集,这些数据集需包含来自相同组织块的整体RNA-seq、单细胞/单核RNA-seq以及空间组学或成像数据。
  • 从相同组织样本生成金标准数据集对于验证和改进去卷积方法至关重要。
  • 未来的去卷积方法必须显式建模细胞大小和总mRNA含量,才能实现准确的细胞类型比例估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。