[论文解读] The p-filter: multi-layer FDR control for grouped hypotheses
p-filter 是一种新颖的多重假设检验程序,通过分层阈值过滤p值,同时控制多个任意非层次化假设划分(如空间、时间或功能分组)下的错误发现率(FDR)。它推广了 Benjamini-Hochberg(BH)程序和 Simes 检验,在不损失统计功效的前提下,提升了发现的精确度,尤其当信号自然分组时效果更佳。
In many practical applications of multiple hypothesis testing using the False Discovery Rate (FDR), the given hypotheses can be naturally partitioned into groups, and one may not only want to control the number of false discoveries (wrongly rejected null hypotheses), but also the number of falsely discovered groups of hypotheses (we say a group is falsely discovered if at least one hypothesis within that group is rejected, when in reality the group contains only nulls). In this paper, we introduce the p-filter, a procedure which unifies and generalizes the standard FDR procedure by Benjamini and Hochberg and global null testing procedure by Simes. We first prove that our proposed method can simultaneously control the overall FDR at the finest level (individual hypotheses treated separately) and the group FDR at coarser levels (when such groups are user-specified). We then generalize the p-filter procedure even further to handle multiple partitions of hypotheses, since that might be natural in many applications. For example, in neuroscience experiments, we may have a hypothesis for every (discretized) location in the brain, and at every (discretized) timepoint: does the stimulus correlate with activity in location x at time t after the stimulus was presented? In this setting, one might want to group hypotheses by location and by time. Importantly, our procedure can handle multiple partitions which are nonhierarchical (i.e. one partition may arrange p-values by voxel, and another partition arranges them by time point; neither one is nested inside the other). We prove that our procedure controls FDR simultaneously across these multiple lay- ers, under assumptions that are standard in the literature: we do not need the hypotheses to be independent, but require a nonnegative dependence condition known as PRDS.
研究动机与目标
- 解决经典 FDR 程序(如 BH)忽略假设中结构化分组的局限性。
- 开发一种方法,可在不依赖层次化划分的前提下,同时在多个分组层次(如空间、时间或功能区域)上控制 FDR。
- 使研究人员能够将领域特定的先验知识(如脑区或基因家族)整合到多重检验中,同时保持统计保障。
- 通过利用分组结构提升发现的精确度,减少错误发现,同时保持高统计功效。
提出的方法
- p-filter 接收 n 个 p 值以及 M ≥ 1 个将这些 p 值划分为组的任意划分。
- 对于每个划分,基于该组中的 p 值计算一个分层特定的阈值,使用广义的逐步提升程序,在用户指定的显著性水平 α_m 下控制每层 m 的 FDR。
- 仅当一个假设在所有层级上均通过 FDR 控制阈值时,才被拒绝(宣布为发现),从而确保多层 FDR 的同步控制。
- 该算法采用递归过滤机制,从最宽松的层级开始,逐步在所有层级上施加更严格的条件。
- 该方法在 PRDS(子集正回归依赖)条件下被证明是有效的,该条件在 fMRI 和基因组数据中通常成立。
- 它推广了 BH(当 M=1 且使用最细粒度划分时)和 Simes 检验(当 M=1 且使用最粗粒度划分时),统一了两种经典程序。
实验结果
研究问题
- RQ1能否在多个非层次化假设划分(如 fMRI 数据中的空间和时间分组)上同时控制 FDR?
- RQ2将领域特定的分组(如脑区或基因家族)纳入分析,是否能相比标准 FDR 程序更精确地发现真实信号?
- RQ3当信号在多个维度上自然分组时,p-filter 与 BH 程序相比在统计功效和精确度方面表现如何?
- RQ4在真实世界的神经科学和基因组学应用中,p-filter 是否能有效减少错误发现,同时保持高统计功效?
- RQ5该方法在大规模多重检验问题中是否具备鲁棒性和可扩展性,尤其适用于复杂、多模态的分组结构?
主要发现
- 在 PRDS 假设下,p-filter 可在所有 M 个划分上实现可证明的同步 FDR 控制,即使这些划分是非层次化的。
- 在模拟实验中,当信号在行和列上分组时,p-filter 的功效几乎与 BH 相当,但错误发现率显著降低。
- 在包含 41,073 × 3 个 p 值的 fMRI 应用中,p-filter 使用 α1=0.05、α2=0.05 和 α3=0.1,成功在个体、时间及空间分组层面上控制了 FDR。
- 与标准 BH 方法相比,该方法减少了错误发现的数量,提升了精确度,且未造成显著的功效损失。
- 在模拟和真实 fMRI 数据中均表明,当信号自然分组时,p-filter 在精确度方面优于 BH。
- 该方法具有高度灵活性和可扩展性,适用于任意数量的任意划分,因此适用于多模态、时空或功能分组,在多种科学领域中均具适用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。