[论文解读] Accumulation tests for FDR control in ordered hypothesis testing
本文提出了一类用于有序先验信息下多重假设检验的累积检验方法,其中假设按其为真实信号的可能性进行排序。新颖的 HingeExp 方法自适应地选择一个发现截断点,以在有限样本中控制修正的 FDR 的同时最大化检验效能,显著优于现有的 Benjamini-Hochberg 和 Storey 方法,无论是在模拟实验还是真实基因表达数据中均表现更优。
Multiple testing problems arising in modern scientific applications can involve simultaneously testing thousands or even millions of hypotheses, with relatively few true signals. In this paper, we consider the multiple testing problem where prior information is available (for instance, from an earlier study under different experimental conditions), that can allow us to test the hypotheses as a ranked list in order to increase the number of discoveries. Given an ordered list of n hypotheses, the aim is to select a data-dependent cutoff k and declare the first k hypotheses to be statistically significant while bounding the false discovery rate (FDR). Generalizing several existing methods, we develop a family of "accumulation tests" to choose a cutoff k that adapts to the amount of signal at the top of the ranked list. We introduce a new method in this family, the HingeExp method, which offers higher power to detect true signals compared to existing techniques. Our theoretical results prove that these methods control a modified FDR on finite samples, and characterize the power of the methods in the family. We apply the tests to simulated data, including a high-dimensional model selection problem for linear regression. We also compare accumulation tests to existing methods for multiple testing on a real data problem of identifying differential gene expression over a dosage gradient.
研究动机与目标
- 解决假设按先验信念排序的多重检验问题,利用该结构以提升发现效能。
- 开发一类通用的累积检验方法,自适应地选择数据依赖的截断点以提升现有顺序检验方法的性能。
- 提出一种新方法 HingeExp,其统计效能高于现有技术,同时保持有限样本的 FDR 控制。
- 通过模拟实验和真实世界应用(如药物剂量梯度下的差异基因表达分析)验证该方法的有效性。
提出的方法
- 该方法使用包含 n 个假设的排序列表,其中先验信息根据其预期信号强度对它们进行排序。
- 提出一类累积检验方法,通过选择数据依赖的截断点 k,将前 k 个假设判定为显著。
- HingeExp 方法使用钩型加权函数,强调早期发现,同时控制修正的 FDR。
- 该方法推广了现有方法如 ForwardStop 和 SeqStep,且在有限样本 FDR 控制方面具有理论保证。
- 该方法可自适应地响应排序列表顶部的信号量,相比非自适应或均匀加权方法,显著提升检验效能。
- 在真实数据分析中使用基于置换的 p 值,对 p 值进行微调以避免 HingeExp 和 ForwardStop 中出现无穷权重。
实验结果
研究问题
- RQ1能否开发一类自适应选择有序假设检验中发现截断点的累积检验方法,同时在有限样本中控制 FDR?
- RQ2HingeExp 方法的效能与现有顺序检验方法(如 SeqStep、ForwardStop 和 Benjamini-Hochberg)相比如何?
- RQ3在高维多重检验问题中,引入先验排序在多大程度上能提升发现效能?
- RQ4在现实数据条件下(包括高维回归和基因表达数据),HingeExp 方法是否能保持 FDR 控制?
- RQ5所提出的框架能否有效应用于真实世界问题,如识别药物剂量梯度下的差异基因表达?
主要发现
- 在一系列目标 FDR 水平下,HingeExp 方法的发现数量显著多于 SeqStep、SeqStep+ 和 ForwardStop,尤其在 α ≤ 0.5 时表现更优。
- 在目标 FDR 水平为 0.5 或更低时,Benjamini-Hochberg 和 Storey 方法仅发现少量假设,而累积检验方法则发现了数千个。
- SeqStep 和 SeqStep+ 方法在性能上几乎无法区分,表明后者微小的校正对效能影响极小。
- 在真实基因表达实验中,累积检验方法通过利用数据的有序结构,显著优于标准 FDR 控制程序。
- 在模拟的高维线性回归和真实差异表达数据中,HingeExp 方法的效能均高于现有方法。
- 理论结果证实,整个累积检验族对修正的 FDR 实现了有限样本控制,且对标准 FDR 实现了渐近控制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。