[Paper Review] Challenges and opportunities to computationally deconvolve heterogeneous tissue with varying cell sizes using single cell RNA-sequencing datasets
This paper identifies that existing deconvolution methods for bulk RNA-seq data fail when applied to tissues with highly variable cell sizes—such as brain or immune tissues—because they conflate cell size and transcriptomic activity, leading to inaccurate cell proportion estimates. The authors advocate for standardized, multi-omics 'gold standard' datasets from the same tissue blocks to enable development and validation of robust, size-aware deconvolution methods.
Deconvolution of cell mixtures in "bulk" transcriptomic samples from homogenate human tissue is important for understanding the pathologies of diseases. However, several experimental and computational challenges remain in developing and implementing transcriptomics-based deconvolution approaches, especially those using a single cell/nuclei RNA-seq reference atlas, which are becoming rapidly available across many tissues. Notably, deconvolution algorithms are frequently developed using samples from tissues with similar cell sizes. However, brain tissue or immune cell populations have cell types with substantially different cell sizes, total mRNA expression, and transcriptional activity. When existing deconvolution approaches are applied to these tissues, these systematic differences in cell sizes and transcriptomic activity confound accurate cell proportion estimates and instead may quantify total mRNA content. Furthermore, there is a lack of standard reference atlases and computational approaches to facilitate integrative analyses, including not only bulk and single cell/nuclei RNA-seq data, but also new data modalities from spatial -omic or imaging approaches. New multi-assay datasets need to be collected with orthogonal data types generated from the same tissue block and the same individual, to serve as a "gold standard" for evaluating new and existing deconvolution methods. Below, we discuss these key challenges and how they can be addressed with the acquisition of new datasets and approaches to analysis.
Motivation & Objective
- To identify and address the critical challenge of cell size heterogeneity in transcriptomic deconvolution of bulk tissue samples.
- To highlight that current deconvolution algorithms, trained on tissues with similar cell sizes, produce biased estimates when applied to tissues like brain or immune systems with extreme size variation.
- To emphasize the lack of standardized reference atlases and integrated multi-omics datasets for validating deconvolution methods.
- To advocate for the generation of orthogonal, multi-modal datasets (e.g., bulk RNA-seq, single-nucleus RNA-seq, spatial omics) from the same tissue blocks to serve as gold standards.
- To promote the development of new computational methods that account for cell size and total mRNA content to improve accuracy in cell proportion inference.
Proposed method
- Propose the creation of multi-omics datasets from the same tissue blocks, integrating bulk RNA-seq, single-cell/nuclei RNA-seq, and spatial-omic or imaging data.
- Use these integrated datasets as 'gold standard' references to benchmark and validate existing and new deconvolution algorithms.
- Develop computational methods that explicitly model cell size and total mRNA content as confounding variables in deconvolution.
- Apply existing deconvolution algorithms to diverse tissues (e.g., brain, immune) to demonstrate systematic bias due to size variation.
- Highlight the need for reference atlases that include not only transcriptomic profiles but also morphometric and subcellular data.
- Encourage the use of orthogonal data modalities to disentangle biological signals from technical artifacts in deconvolution.
Experimental results
Research questions
- RQ1How do differences in cell size and total mRNA content affect the accuracy of current deconvolution algorithms in bulk RNA-seq analysis?
- RQ2To what extent do existing deconvolution methods conflate cell size with cell type proportions when applied to heterogeneous tissues?
- RQ3What are the key limitations of current single-cell/nuclei RNA-seq reference atlases in enabling accurate deconvolution of tissues with extreme cell size variation?
- RQ4How can multi-omics datasets from the same tissue block serve as gold standards for evaluating deconvolution methods?
- RQ5What computational frameworks are needed to disentangle cell size effects from true biological cell type composition in deconvolution?
Key findings
- Existing deconvolution methods systematically misestimate cell proportions in tissues with highly variable cell sizes, such as brain and immune tissues.
- The bias arises because algorithms conflate differences in total mRNA content due to cell size with actual cell type abundance.
- Current single-cell/nuclei RNA-seq reference atlases are insufficient for accurate deconvolution in tissues with extreme cell size heterogeneity.
- There is a critical lack of standardized, multi-modal reference datasets that include bulk RNA-seq, single-cell/nuclei RNA-seq, and spatial-omic or imaging data from the same tissue blocks.
- The development of gold standard datasets from the same tissue samples is essential for validating and improving deconvolution methods.
- Future deconvolution methods must explicitly model cell size and total mRNA content to achieve accurate cell type proportion estimation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.