Skip to main content
QUICK REVIEW

[论文解读] Testing hypotheses on a tree: new error rates and controlling strategies

Marina Bogomolov, Christine B. Peterson|arXiv (Cornell University)|May 22, 2017
Genetic Associations and Epidemiology参考文献 23被引用 16
一句话总结

本文提出 TreeBH,一种新颖的多重假设检验方法,可在层次树结构的假设体系中,于多个分辨率层级上控制错误发现率(FDR)。通过利用依赖性假设与顺序检验算法,TreeBH 在保持 FDR 控制的同时,相较于现有方法提升了统计功效,该优势在模拟实验及 GTEx 和微生物组数据的真实应用中得到验证。

ABSTRACT

We introduce a multiple testing procedure (TreeBH) which addresses the challenge of controlling error rates at multiple levels of resolution. Conceptually, we frame this problem as the selection of hypotheses which are organized hierarchically in a tree structure. We describe a fast algorithm for the proposed sequential procedure, and prove that it controls relevant error rates given certain assumptions on the dependence among the p-values. Through simulations, we demonstrate that TreeBH offers the desired guarantees under a range of dependency structures (including one similar to that encountered in genome-wide association studies) and that it has the potential of gaining power over alternative methods. We also introduce a modified version of TreeBH which we prove to control the relevant error rates under any dependency structure. We conclude with two case studies: we first analyze data collected as part of the Genotype-Tissue Expression (GTEx) project, which aims to characterize the genetic regulation of gene expression across multiple tissues in the human body, and secondly, data examining the relationship between the gut microbiome and colorectal cancer.

研究动机与目标

  • 解决在层次假设检验中于多个分辨率层级控制错误率的挑战。
  • 开发一种在假设以树形结构组织时仍能保持 FDR 控制的方法,而非将所有假设同等对待。
  • 通过利用假设的层次结构提升统计功效,减少对细粒度假设的无谓检验。
  • 提供一种适用于顺序数据收集的框架,其中假设按分辨率顺序进行检验。
  • 确保在各类 p 值依赖结构下具有鲁棒性,包括真实基因组数据中常见的依赖结构。

提出的方法

  • 提出 TreeBH,一种按树结构自顶向下顺序进行假设检验的多重检验程序。
  • 采用逐步程序,仅当某节点的 p 值低于基于其下游节点动态调整的阈值时,才拒绝该节点的假设。
  • 在树结构上满足正回归依赖性(PRDS)的条件下,控制树中每一层级的 FDR。
  • 提出一种改进的 TreeBH 变体,可在任意依赖结构下控制 FDR,提升方法的鲁棒性。
  • 实现一种快速算法,按后序遍历树结构,从叶节点到根节点递归计算临界值。
  • 通过在 GTEx 和微生物组研究中将 SNPs(第 1 层)、基因(第 2 层)和组织(第 3 层)组织成层次树结构,将该方法应用于真实数据。

实验结果

研究问题

  • RQ1是否可以在树状假设空间中,同时在多个分辨率层级上控制 FDR?
  • RQ2利用层次结构是否能相比标准 FDR 控制方法(如 BH 或 BB)提升统计功效?
  • RQ3在真实依赖结构(如多组织 eQTL 研究中的结构)下,TreeBH 的表现如何?
  • RQ4该方法能否在不依赖强假设的前提下,适用于 p 值间任意依赖关系?
  • RQ5当应用于 GTEx 和微生物组研究等真实世界多组学数据集时,TreeBH 是否能保持错误率控制?

主要发现

  • 在涵盖多种依赖结构(包括模拟真实多组织 eQTL 数据的结构)的条件下,TreeBH 在树的所有层级上近似控制了目标 FDR。
  • TreeBH 在 FDR、sFDR 和统计功效方面与基准方法 BB 相当,但在稀疏信号条件下错误率控制更优。
  • 在包含 8,713 个 SNPs 和 250 个基因、共 5 个组织的模拟实验中,TreeBH 在 FDR 控制方面优于 BH 方法,后者未能维持目标错误率。
  • 改进的 TreeBH 变体在任意依赖结构下成功控制了 FDR,将方法的鲁棒性扩展至超越 PRDS 假设的范围。
  • 在 GTEx 案例研究中,TreeBH 在多个分辨率层级上识别出具有生物学意义的 eGenes 和 eSNPs,且错误率得到良好控制。
  • 在微生物组研究中,TreeBH 有效检测到属和科层级的关联,展示了其在复杂微生物群落分析中的实用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。