Skip to main content
QUICK REVIEW

[论文解读] False discovery rate control for identifying simultaneous signals

Sihai Dave Zhao|arXiv (Cornell University)|Dec 14, 2015
Genetic Mapping and Diversity in Plants and Animals参考文献 52被引用 3
一句话总结

本文提出了一种无需调节参数的非参数方法,用于识别在两个或更多独立研究中同时显著的信号(即特征),并可证明地控制错误发现率(FDR)。该方法无需了解原假设或备择假设的分布,且在模拟实验和精神疾病全基因组关联研究(GWAS)分析中,相较于现有方法展现出更高的统计功效和更优的错误控制能力。

ABSTRACT

It is frequently of interest to jointly analyze multiple sequences of multiple tests in order to identify simultaneous signals, defined as features tested in two or more independent studies that are significant in each. For example, researchers often wish to discover genetic variants that are significantly associated with multiple traits. This paper proposes a false discovery rate control procedure for identifying simultaneous signals in two studies. A pair of test statistics is available for each feature, and the goal is to identify features for which both are non-null. Error control is difficult due to the composite nature of a non-discovery, as one of the tests in the pair can still be non-null. Very few existing methods have high power while still provably controlling the false discovery rate. This paper proposes a simple, fast, tuning parameter-free nonparametric procedure that can be shown to provide asymptotically conservative false discovery rate control. Surprisingly, the procedure does not require knowledge of either the null or the alternative distributions of the test statistics. In simulations, the proposed method had higher power and better error control than existing procedures. In an analysis of genome-wide association study results from five psychiatric disorders, it identified more pairs of disorders that share simultaneously significant genetic variants, as well as more variants themselves, compared to other methods. The proposed method is available in the R package ssa.

研究动机与目标

  • 解决在控制错误发现率的同时,识别在多个独立研究中均显著的特征的挑战。
  • 开发一种方法,在复合型非发现(即成对中的一个检验仍可能为非零)情形下,仍能保持强错误控制。
  • 创建一种计算快速、无需调节参数且无需了解检验统计量原假设或备择分布的程序。
  • 相较于现有FDR控制方法,提升检测同步信号的统计功效。
  • 实现在全基因组关联研究中对多种精神疾病共享遗传变异的稳健检测。

提出的方法

  • 该方法基于两个研究中成对检验统计量的联合分布,采用非参数方法,不假设原假设或备择分布的参数形式。
  • 采用逐步提升程序,根据基于检验统计量的准则对特征进行排序,以渐近保守地控制FDR。
  • 该程序对检验统计量之间的依赖关系具有鲁棒性,且无需估计效应大小或方差成分。
  • 通过利用联合p值或检验统计量的排序,识别出在两个研究中均拒绝原假设的特征,从而在弱依赖假设下实现FDR控制。
  • 该方法已通过R包 ssa 实现,便于在多研究基因组数据中应用。
  • 该方法通过从检验统计量的经验分布中推导出的数据驱动阈值规则,避免了调节参数的使用。

实验结果

研究问题

  • RQ1在识别两个研究中的同步信号时,一种非参数且无需调节参数的方法能否有效控制错误发现率?
  • RQ2与现有用于成对研究中多重检验的FDR控制程序相比,该方法在功效和错误控制方面表现如何?
  • RQ3该方法是否在无需了解检验统计量原假设或备择分布的情况下,仍能保持FDR控制?
  • RQ4在真实世界的全基因组关联研究数据中,该方法在多大程度上能检测到多种精神疾病之间的共享遗传变异?
  • RQ5该方法在成对研究中对检验统计量之间的依赖关系是否具有鲁棒性?

主要发现

  • 在模拟实验中,该方法相较于现有程序展现出更高的统计功效,尤其在中等到高强度信号的情境下。
  • 该方法实现了渐近保守的FDR控制,即实际FDR未超过名义水平,即使在复杂的依赖结构下也成立。
  • 在一项涉及五种精神疾病的GWAS分析中,该方法识别出的共享显著遗传变异的疾病对数量多于其他方法。
  • 该方法检测到更多在多个研究中均显著的个体遗传变异,表明其敏感性更高。
  • 实现该方法的R包 ssa 使得多研究基因组数据的高效且便捷应用成为可能。
  • 该方法在类型I错误控制和功效方面均优于现有方法,且无需依赖分布假设。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。