[论文解读] Adapting BH to One- and Two-Way Classified Structures of Hypotheses
本文提出了一种数据自适应程序,用于在单向和双向分类的假设结构下控制多重检验中的假发现率(FDR),通过使用加权p值扩展了Benjamini-Hochberg(BH)程序。该文提出了一个Oracle方法和数据自适应方法,利用基于结构分类的组特异性权重,在PRDS条件下实现了非渐近FDR控制,并在现有方法之上提升了统计功效。
Multiple testing literature contains ample research on controlling false discoveries for hypotheses classified according to one criterion, which we refer to as one-way classified hypotheses. Although simultaneous classification of hypotheses according to two different criteria, resulting in two-way classified hypotheses, do often occur in scientific studies, no such research has taken place yet, as far as we know, under this structure. This article produces procedures, both in their oracle and data-adaptive forms, for controlling the overall false discovery rate (FDR) across all hypotheses effectively capturing the underlying one- or two-way classification structure. They have been obtained by using results associated with weighted Benjamini-Hochberg (BH) procedure in their more general forms providing guidance on how to adapt the original BH procedure to the underlying one- or two-way classification structure through an appropriate choice of the weights. The FDR is maintained non-asymptotically by our proposed procedures in their oracle forms under positive regression dependence on subset of null $p$-values (PRDS) and in their data-adaptive forms under independence of the $p$-values. Possible control of FDR for our data-adaptive procedures in certain scenarios involving dependent $p$-values have been investigated through simulations. The fact that our suggested procedures can be superior to contemporary practices has been demonstrated through their applications in simulated scenarios and to real-life data sets. While the procedures proposed here for two-way classified hypotheses are new, the data-adaptive procedure obtained for one-way classified hypotheses is alternative to and often more powerful than those proposed in Hu et al. (2010).
研究动机与目标
- 为解决现有多重检验程序在有效整合假设的一维或二维分类结构方面的不足。
- 开发在保持强误差率控制的同时利用组级别信息的FDR控制程序。
- 将加权BH程序扩展至二维分类假设,填补了多重检验文献中此前未被研究的设置。
- 通过最优权重分配整合结构信息,提升统计功效。
- 提供基于数据估计权重的数据自适应程序,避免依赖外部知识。
提出的方法
- 使用基于单向或双向分类结构中组成员关系的加权Benjamini-Hochberg(BH)程序。
- 对加权p值应用逐步提升程序,临界常数与权重成比例,确保在Oracle版本下满足PRDS条件的FDR控制。
- 采用Storey型估计方法估计原假设比例,以构建数据自适应权重,实现实际应用。
- 针对二维分类,提出使用行权重和列权重(基于边际组大小)的特定权重组合。
- 通过条件期望和PRDS下的随机单调性推导理论FDR上界,证明了非渐近控制。
- 通过模拟和真实数据应用验证性能,与Hu等(2010)和Ignatiadis等(2016)等现有方法进行比较。
实验结果
研究问题
- RQ1如何将Benjamini-Hochberg程序调整为在基于组大小的组特异性权重下,保持单向分类假设结构中的FDR控制?
- RQ2在二维分类假设结构中,如何最优地分配权重以在控制FDR的同时最大化统计功效?
- RQ3基于估计原假设比例的数据自适应程序是否能在FDR控制和功效方面优于Oracle方法和现有方法?
- RQ4在p值存在依赖关系时,所提出的方法表现如何?在何种条件下可确保FDR控制?
- RQ5在真实世界和模拟的多重检验场景中,所提出程序相较于当前实践可实现多大程度的改进?
主要发现
- 所提出的Oracle程序在p值满足PRDS条件时,即使权重基于组大小选择,也能实现非渐近FDR控制。
- 数据自适应的一维分组BH程序被证明比Hu等(2010)提出的方法更具功效,尤其在组大小差异显著时。
- 模拟结果表明,二维分组BH程序在独立性条件下保持了FDR控制,并在结构化设置中表现出更高的功效。
- 数据自适应的二维程序优于忽略一个分类维度的单向方法的朴素应用。
- 理论结果证实,加权逐步提升程序下的FDR被个体原假设拒绝概率之和所上界控制,从而确保了控制效果。
- 真实数据集的实证应用证实,该方法在识别显著假设方面优于现有方法,同时保持了FDR控制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。