[论文解读] A New Approach for Large Scale Multiple Testing with Application to FDR Control for Graphically Structured Hypotheses
本文提出了一种用于大规模多重检验的新型通用框架,可实现具有有向无环图(DAG)结构的假设的错误发现率(FDR)控制程序。通过将FDR控制问题转化为每家族误差率(PFER)控制问题,该方法确保被拒绝的假设保留原始DAG结构,相较于标准程序(如BH),显著提升了统计功效与可解释性。
In many large scale multiple testing applications, the hypotheses often have a known graphical structure, such as gene ontology in gene expression data. Exploiting this graphical structure in multiple testing procedures can improve power as well as aid in interpretation. However, incorporating the structure into large scale testing procedures and proving that an error rate, such as the false discovery rate (FDR), is controlled can be challenging. In this paper, we introduce a new general approach for large scale multiple testing, which can aid in developing new procedures under various settings with proven control of desired error rates. This approach is particularly useful for developing FDR controlling procedures, which is simplified as the problem of developing per-family error rate (PFER) controlling procedures. Specifically, for testing hypotheses with a directed acyclic graph (DAG) structure, by using the general approach, under the assumption of independence, we first develop a specific PFER controlling procedure and based on this procedure, then develop a new FDR controlling procedure, which can preserve the desired DAG structure among the rejected hypotheses. Through a small simulation study and a real data analysis, we illustrate nice performance of the proposed FDR controlling procedure for DAG-structured hypotheses.
研究动机与目标
- 解决在已知图形(DAG)结构的假设下,大规模多重检验中错误发现率(FDR)控制的挑战。
- 提出一种通用的方法论框架,通过将FDR控制程序的构建简化为PFER控制程序的设计,从而简化新多重检验程序的开发。
- 确保被拒绝的假设集合保持原始DAG结构,以保留层次完整性,提升可解释性。
- 在独立性假设下,通过实证比较证明该方法相较于标准程序(如BH)具有更高的统计功效,同时保持强误差率控制。
- 为所提出的框架提供理论保证,特别是在DAG结构化假设下的FDR控制。
提出的方法
- 提出一种通用方法(定理3.1),将FDR控制问题简化为PFER控制问题,从而简化新型多重检验程序的设计。
- 在独立性假设下,为DAG结构化假设设计一种具体的PFER控制程序(定理4.1),采用基于父子关系的递归加权方案。
- 通过转换PFER控制程序,构建一种新的FDR控制程序(定理5.1),确保仅当一个假设的所有父假设均被拒绝时,该假设才被拒绝。
- 使用涉及复合泊松分布生存函数逆函数的临界值函数,以校准拒绝阈值,同时保持DAG结构。
- 应用一种递归算法,按拓扑顺序处理假设,从叶节点开始,逐步向根节点推进,以计算临界值和拒绝决策。
- 利用所推导程序的PFER控制性质,通过耦合论证和指示变量界,证明在独立性假设下FDR可被控制。
实验结果
研究问题
- RQ1能否开发一种通用框架,以简化结构化假设下FDR控制程序的构建?
- RQ2如何利用每家族误差率(PFER)控制框架,推导出适用于DAG结构化假设的有效FDR控制程序?
- RQ3与标准FDR程序(如BH)相比,引入DAG结构在多大程度上提升了统计功效?
- RQ4所提出的方法能否确保被拒绝假设的集合构成一个有效的子DAG,从而保留原始层次结构?
- RQ5在所提方法下,特别是独立性假设下,FDR控制能提供哪些理论保证?
主要发现
- 所提出的FDR控制程序在被拒绝的假设中保持了DAG结构,确保若一个假设被拒绝,则其在DAG中的所有祖先假设也均被拒绝。
- 在模拟研究和真实数据分析中,该方法的统计功效均高于Benjamini-Hochberg(BH)程序,尤其当底层DAG结构具有信息量时更为显著。
- 理论分析表明,在独立性假设下FDR可被控制,其证明依赖于通过PFER控制框架对假发现数期望的界进行控制。
- 结果显示,BH程序的修改版本可保留DAG结构,但以牺牲统计功效为代价,表明结构保留与效率之间存在权衡。
- 临界值函数通过基于后代叶节点数量的递归加权方案推导得出,该方法在适应图拓扑结构的同时,确保了适当的错误率控制。
- 该框架具有通用性,只要存在相应的PFER控制程序,即可扩展至其他误差率类型。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。