Skip to main content
QUICK REVIEW

[论文解读] Finding the centre: corrections for asymmetry in high-throughput sequencing datasets

Jia Wu, Jean M. Macklaim|arXiv (Cornell University)|Apr 6, 2017
Geochemistry and Geologic Mapping被引用 9
一句话总结

本文通过扩展对数比变换(特别是CLR和ALR),针对高通量测序(HTS)数据中的非对称性和组合偏差,提出了一种稳健的贝叶斯组合分析方法。研究证明,未经校正的非对称性会导致差异表达分析中出现假阳性与假阴性结果,并提出基于预设参考的校正方法,该方法已集成至Bioconductor的ALDEx2工具中,在稀疏且不平衡的数据集中显著提升了推断准确性。

ABSTRACT

High throughput sequencing is a technology that allows for the generation of millions of reads of genomic data regarding a study of interest, and data from high throughput sequencing platforms are usually count compositions. Subsequent analysis of such data can yield information on tran- scription profiles, microbial diversity, or even relative cellular abundance in culture. Because of the high cost of acquisition, the data are usually sparse, and always contain far fewer observations than variables. However, an under-appreciated pathology of these data are their often unbalanced nature: i.e, there is often systematic variation between groups simply due to presence or absence of features, and this variation is important to the biological interpretation of the data. A simple example would be comparing transcriptomes of yeast cells with and without a gene knockout. This causes samples in the comparison groups to exhibit widely varying centres. This work extends a previously described log-ratio transformation method that allows for variable comparisons between samples in a Bayesian compositional context. We demonstrate the pathology in modelled and real unbalanced experimental designs to show how this dramatically causes both false negative and false positive inference. We then introduce several approaches to demonstrate how the pathologies can be addressed. An extreme example is presented where only the use of a predefined basis is appropriate. The transformations are implemented as an extension to a general compositional data analysis tool known as ALDEx2 which is available on Bioconductor.

研究动机与目标

  • 解决高通量测序(HTS)数据中被低估的非对称性病理问题,该问题源于特征存在/缺失引发的系统性变异,而非真实的生物学变化。
  • 证明基于计数的标准方法因组合偏差而失效,导致差异表达分析中出现假阳性与假阴性结果。
  • 通过引入基于预设基底的非对称性校正,扩展贝叶斯组合数据分析方法,提升在稀疏且不平衡实验设计中的鲁棒性。
  • 通过集成至ALDEx2(一个Bioconductor工具)提供实用解决方案,实现对HTS数据的准确且可靠的解读。

提出的方法

  • 使用中心对数比(CLR)变换,定义为 log(xi / g(x)),其中 g(x) 为所有特征的几何平均值,以稳定比值并确保尺度不变性。
  • 应用加法对数比(ALR)变换,log(xi / xD),选择一个选定的参考特征(如管家基因或稳定OTU)作为分母,以实现相对比较。
  • 通过选择预设且稳定的特征作为参考,引入基于基底的校正方法,尤其在数据高度不平衡或稀疏时至关重要。
  • 在组合分析框架内采用贝叶斯估计,以提升在高维、稀疏HTS数据中的鲁棒性,并减少错误推断。
  • 使用模拟数据和真实HTS数据集(包括阴道微生物组研究和酿酒酵母基因敲除转录组)验证方法的有效性。
  • 将该方法作为ALDEx2的扩展实现,ALDEx2是广泛用于HTS数据组合分析的Bioconductor软件包。

实验结果

研究问题

  • RQ1在真实实验设计中,非平衡HTS数据导致的组合偏差在差异表达推断中会产生何种影响?
  • RQ2标准计数方法(如负二项分布模型)在多大程度上因HTS数据固有的组合性质而失效?
  • RQ3在非对称数据集中,ALR变换中预设基底的选择能否显著降低假阳性与假阴性率?
  • RQ4在稀疏且非平衡的HTS数据中,参考特征的选择如何影响差异表达结果的鲁棒性?
  • RQ5结合CLR与ALR校正的贝叶斯组合建模能否在统计推断上优于传统归一化与检验方法?

主要发现

  • 由于特征存在/缺失导致的非平衡HTS数据,若使用标准计数模型分析,将导致显著的假阳性与假阴性推断。
  • CLR变换能有效稳定比值并保持尺度不变性,但在样本总读数显著变化时会失效。
  • 使用预设且稳定的参考特征(如管家基因或丰度高的OTU)的ALR变换,在高度非对称数据集中提供了稳健的解决方案。
  • 在极端情况下(如仅一个特征始终存在),只有预设基底(ALR)能提供可靠推断,而CLR因几何均值不稳定性而失效。
  • 所提出的校正方法在真实数据集中显著提升了差异表达的准确性,包括酿酒酵母基因敲除研究和人类阴道微生物组队列研究。
  • 集成至ALDEx2使研究人员可直接应用这些校正方法,从而提升组合HTS数据分析的可重复性与可靠性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。