[论文解读] On Online Control of False Discovery Rate
本文提出了LORD和LORD兩種線上程序,用於在序列多重假設檢驗中控制偽發現在率(FDR)與邊際偽發現在率(mFDR),其中決策需根據過去資訊即時作出。這些方法在任意依賴結構下確保FDR控制,且在具有稀疏替代假設的混合模型下實現近乎線性的發現速率。
Multiple hypotheses testing is a core problem in statistical inference and arises in almost every scientific field. Given a sequence of null hypotheses $\mathcal{H}(n) = (H_1,..., H_n)$, Benjamini and Hochberg \cite{benjamini1995controlling} introduced the false discovery rate (FDR) criterion, which is the expected proportion of false positives among rejected null hypotheses, and proposed a testing procedure that controls FDR below a pre-assigned significance level. They also proposed a different criterion, called mFDR, which does not control a property of the realized set of tests; rather it controls the ratio of expected number of false discoveries to the expected number of discoveries. In this paper, we propose two procedures for multiple hypotheses testing that we will call "LOND" and "LORD". These procedures control FDR and mFDR in an \emph{online manner}. Concretely, we consider an ordered --possibly infinite-- sequence of null hypotheses $\mathcal{H} = (H_1,H_2,H_3,...)$ where, at each step $i$, the statistician must decide whether to reject hypothesis $H_i$ having access only to the previous decisions. To the best of our knowledge, our work is the first that controls FDR in this setting. This model was introduced by Foster and Stine \cite{alpha-investing} whose alpha-investing rule only controls mFDR in online manner. In order to compare different procedures, we develop lower bounds on the total discovery rate under the mixture model and prove that both LOND and LORD have nearly linear number of discoveries. We further propose adjustment to LOND to address arbitrary correlation among the $p$-values. Finally, we evaluate the performance of our procedures on both synthetic and real data comparing them with alpha-investing rule, Benjamin-Hochberg method and a Bonferroni procedure.
研究动机与目标
- 解決在線上、序列多重假設檢驗中控制FDR的挑戰,其中假設數量可能無限,且決策僅能基於過去資訊作出。
- 發展在線上環境下控制FDR與mFDR的程序,克服現有方法(如alpha-investing與Benjamini-Hochberg)的限制。
- 在具有稀疏非零假設的混合模型下,建立發現速率的理論保障,並允許p值之間任意依賴。
- 提出LORD的一種強健變體,以處理p值之間的任意相關性。
- 透過合成資料與真實資料,實證評估其性能,對比既有的方法如Benjamini-Hochberg、alpha-investing與Bonferroni。
提出的方法
- 提出LORD(Lond)與LORD(Lord),分別為控制FDR與mFDR的線上程序,利用基於過去發現的自適應臨界值。
- 採用遞迴更新規則來動態調整顯著性臨界值,其依據為過去發現的數量與衰減參數。
- 在混合模型下,利用更新過程框架推導出預期發現數量的理論邊界。
- 提出LORD的一種改良版本,引入修正項以處理p值之間的任意依賴。
- 運用隨機 dominance 論證與更新理論,證明預期的發現間隔時間有界,從而確保線性發現速率。
- 應用遞迴不等式與數學歸納法證明,臨界值與其期望路徑的偏離量在時間上保持有界。
实验结果
研究问题
- RQ1在線上、序列多重檢驗架構中,是否能控制FDR,其中假設逐一到達,且決策僅能基於過去資訊?
- RQ2在未知總假設數量的情況下,發現力與FDR控制之間的最優權衡為何?
- RQ3在混合模型下,當真實非零假設的稀疏性變化時,線上FDR程序的發現速率如何變化?
- RQ4FDR控制是否能在p值之間任意依賴的情況下維持?如何修正此類依賴?
- RQ5LORD與LORD在實證表現上與經典方法(如Benjamini-Hochberg與alpha-investing)相比如何?
主要发现
- LORD與LORD在混合模型下實現近乎線性的發現速率,其中每個假設以固定機率ε為非零假設,且在非零假設下p值獨立同分佈。
- 預期發現數量隨檢驗假設數量線性增長,其增長速率趨近於1/μ,其中μ為平均發現間隔時間。
- 這些程序在線上環境下控制FDR與mFDR,且透過改良版LORD在任意依賴結構下證明FDR控制。
- 理論分析顯示,臨界值與其期望路徑的偏離量保持有界,確保系統穩定與長期控制。
- 在合成資料與真實資料上的實證評估顯示,LORD與LORD在發現力方面優於alpha-investing、Benjamini-Hochberg與Bonferroni,同時維持FDR控制。
- 該方法具備強健的理論保障,包括對預期偽發現數量的邊界控制,即使總假設數量未知或無限亦成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。