Skip to main content
QUICK REVIEW

[论文解读] A Practitioner's Guide to Multiple Testing Error Rates

Jonathan D. Rosenblatt|arXiv (Cornell University)|Apr 17, 2013
Statistical Methods in Clinical Trials参考文献 26被引用 4
一句话总结

本文为研究人员在多重假设检验中选择合适的错误率提供了实用指南,比较了家族错误率(FWER)、错误发现率(FDR)及相关度量。通过遗传学、医学影像、心理学和教育政策等真实案例,展示了如何根据研究背景在这些错误率之间进行选择,强调在大规模检验场景中,FDR控制通常比FWER更具统计功效且更合适。

ABSTRACT

It is quite common in modern research, for a researcher to test many hypotheses. The statistical (frequentist) hypothesis testing framework, does not scale with the number of hypotheses in the sense that naively performing many hypothesis tests will probably yield many false findings. Indeed, statistical "significance" is evidence for the presence of a signal within the noise expected in a single test, not in a multitude. In order to protect himself from an uncontrolled number of erroneous findings, a researcher has to consider of the type or errors he wishes to avoid and select the adequate procedure for that particular error type and data structure. A quick search of the tag [multiple-comparisons] in the statistics Questions & Answers web site Cross Validates (http://stats.stackexchange.com) demonstrates the amount of confusion this task can actually cause. This was also a point made at the 2009 Multiple Comparisons conference in Tokyo. In an attempt to offer guidance, we review possible error types for multiple testing, and demonstrate them with some practical examples, which clarify the formalism. Finally, we include some notes on the software implementations of the methods discussed. The emphasis of this manuscript is on the error-rates, and not on the procedures themselves. We do try to name several procedures in this manuscript where appropriate. P-value adjustment will not be discussed as it is procedure specific. I.e., it is the choice of a procedure that defines the p-value adjustment, and not the error rate itself. Simultaneous confidence intervals will, also, not be discussed.

研究动机与目标

  • 阐明多重假设检验中常用错误率——FWER与FDR之间的区别。
  • 指导研究人员根据其科学目标和数据背景选择最合适的错误率。
  • 通过遗传学、医学影像、心理学和教育政策等真实案例,展示错误率选择的实际影响。
  • 强调在大规模检验中,尤其当重视结果复现与发现时,FDR控制通常比FWER控制更具统计功效且更合适。
  • 澄清常见误解,例如将Benjamini-Hochberg程序等同于FDR本身,并阐明局部fdr与后验概率在决策中的作用。

提出的方法

  • 回顾并定义关键错误率:家族错误率(FWER)、错误发现率(FDR)及其弱控制与强控制变体。
  • 引入局部错误发现率(local fdr)的概念,即给定检验统计量时,原假设为真的后验概率,用作决策标准。
  • 解释FDR定义为所有发现中假发现比例的期望值:E[V/R],当R=0时FDP=0。
  • 建议将局部fdr ≤ 0.2作为实际的拒绝阈值,该阈值在特定条件下近似实现边际FDR控制在约0.1水平。
  • 讨论优化框架,如在FDR ≤ α条件下最小化FNR,或在mFDR ≤ α条件下最小化mFNR,以提升统计功效。
  • 强调基于局部fdr的程序依赖于从检验统计量分布中估计原假设概率,且假设非零效应存在聚集现象。

实验结果

研究问题

  • RQ1FWER与FDR的关键区别是什么?在多重检验场景中,各自应在何时使用?
  • RQ2研究人员如何根据其科学目标(如最小化假阳性与最大化真发现)选择合适的错误率?
  • RQ3局部fdr在大规模假设检验中如何提升统计功效与可解释性?
  • RQ4Benjamini-Hochberg程序与FDR之间有何关系?为何将两者等同是错误的?
  • RQ5在真实应用(如遗传关联研究)中,使用不同错误率阈值(如α = 0.05与α = 0.1)会产生何种影响?

主要发现

  • FWER控制更为保守,当即使一个假发现也无法接受时(如药物注册试验)更为合适。
  • FDR控制在大规模检验中更具统计功效且更适用,例如在基因组学中,可接受少量假发现。
  • 使用局部fdr阈值0.2近似实现边际FDR控制在0.1水平,提供了一种实用且可解释的决策规则。
  • Benjamini-Hochberg程序是FDR控制的一种方法,但不能与FDR本身等同,这是常见的误解。
  • 在已知效应大小的场景中(如预期z值接近7),忽略先验知识会导致检验功效不足及次优的拒绝规则。
  • 基于局部fdr的程序可通过利用检验统计量的经验分布提升功效,尤其在非零效应聚集时更为显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。