Skip to main content
QUICK REVIEW

[论文解读] Generalized Statistical Tests for mRNA and Protein Subcellular Spatial Patterning against Complete Spatial Randomness

Jonathan Warrell, Anca F. Savulescu|arXiv (Cornell University)|Feb 20, 2016
Statistical Methods and Inference参考文献 26被引用 4
一句话总结

本文提出了一类广义的统计估计量,用于空间统计分析——如Ripley的K、H和L函数、聚类指数及聚类程度——使mRNA和蛋白质亚细胞定位模式能够相对于任意随机测度下的完全空间随机性(CSR)进行分析,而不仅限于点过程。该方法通过基于卷积的K估计量将这些统计量扩展至连续强度数据(如蛋白质荧光信号),并结合基于置换的零假设检验,成功揭示了极化小鼠胚胎成纤维细胞中mRNA与蛋白质的聚类相关性,包括此前未报告的mRNA定位模式。

ABSTRACT

We derive generalized estimators for a number of spatial statistics that have been used in the analysis of spatially resolved omics data, such as Ripley's K, H and L functions, clustering index, and degree of clustering, which allow these statistics to be calculated on data modelled by arbitrary random measures (RMs). Our estimators generalize those typically used to calculate these statistics on point process data, allowing them to be calculated on RMs which assign continuous values to spatial regions, for instance to model protein intensity. The clustering index (H*) compares Ripley's H function calculated empirically to its distribution under complete spatial randomness (CSR), leading us to consider CSR null hypotheses for RMs which are not point-processes when generalizing this statistic. We thus consider restricted classes of completely random measures which can be simulated directly (Gamma processes and Marked Poisson Processes), as well as the general class of all CSR RMs, for which we derive an exact permutation-based H* estimator. We establish several properties of the estimators, including bounds on the accuracy of our general Ripley K estimator, its relationship to a previous estimator for the cross-correlation measure, and the relationship of our generalized H* estimator to previous statistics. To test the ability of our approach to identify spatial patterning, we use Fluorescent In Situ Hybridization (FISH) and Immunofluorescence (IF) data to probe for mRNA and protein subcellular localization patterns respectively in polarizing mouse fibroblasts on micropattened cells. We observe correlated patterns of clustering over time for corresponding mRNAs and proteins, suggesting a deterministic effect of mRNA localization on protein localization for several pairs tested, including one case in which spatial patterning at the mRNA level has not been previously demonstrated.

研究动机与目标

  • 将Ripley的K、H和L函数等空间统计量从点过程推广至任意随机测度,包括连续强度数据(如蛋白质荧光信号)。
  • 为非点过程随机测度(包括伽马过程和标记泊松过程)定义并计算有效的CSR零假设。
  • 开发一种广义聚类指数(H*)与聚类程度估计量,可同时适用于离散与连续空间数据。
  • 检验该方法在真实组学数据中检测生物相关空间模式的能力,特别是mRNA与蛋白质的亚细胞定位模式。
  • 证明mRNA定位在极化细胞中驱动确定性的蛋白质定位模式,即使在mRNA模式此前未被检测到的情况下亦成立。

提出的方法

  • 提出一种基于卷积的广义Ripley K函数估计量,适用于任意随机测度,包括连续值空间数据。
  • 推导出在广义CSR随机测度类下,广义H*聚类指数的精确置换检验估计量,避免依赖参数假设。
  • 使用受限的CSR随机测度类(如伽马过程和标记泊松过程)进行模拟,以估计临界分位数,适用于模型拟合。
  • 应用EM算法将标记泊松过程拟合至观测数据,从而在这些特定零模型下实现CSR的模拟。
  • 整合FISH(mRNA)与IF(蛋白质)的荧光显微镜数据,映射微图案化NIH/3T3成纤维细胞中的亚细胞空间分布。
  • 从微管蛋白IF的z-序列构建3D细胞高度图,以实现空间归一化,并在3D中实现精确的空间分析。

实验结果

研究问题

  • RQ1Ripley的K、H和L函数能否被广义化至连续随机测度(如蛋白质强度分布)?是否可超越传统点过程假设?
  • RQ2当零假设模型并非简单泊松过程时,如何在CSR零假设下计算聚类统计量的有效临界分位数?
  • RQ3在极化细胞中,mRNA与蛋白质的空间模式随时间的关联程度如何?这种关联是否可归因于mRNA的确定性定位?
  • RQ4所提出的广义聚类指数能否检测到先前使用标准点过程方法未发现的mRNA分布空间模式?
  • RQ5在成纤维细胞中观察到的mRNA与蛋白质定位之间的空间相关性,是否暗示mRNA定位在决定蛋白质亚细胞定位中起因果作用?

主要发现

  • 广义K估计量在精度上具有界,且在交叉相关测度下与先前估计量等价,验证了其一致性。
  • 所提出的基于置换的H*估计量可在所有CSR随机测度类下提供精确的CSR检验,实现无需参数假设的稳健统计推断。
  • 该方法成功检测到在多个时间点(2–7小时)的微图案化小鼠成纤维细胞中,mRNA(通过FISH)与蛋白质(通过IF)分布的空间聚类。
  • 在时间序列上观察到相应mRNA与蛋白质聚类模式之间存在显著相关性,表明mRNA定位对蛋白质定位具有确定性影响。
  • 首次在一种此前未知具有此类组织结构的基因上识别出mRNA水平的空间模式,凸显该方法的高灵敏度。
  • 该方法揭示mRNA定位是蛋白质亚细胞定位的关键驱动因素,并获得广义空间统计的强有力支持。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。