[论文解读] Asymptotic Bayes optimality under sparsity for generally distributed effect sizes under the alternative
本文将稀疏性下的渐近贝叶斯最优性(ABOS)扩展至一般分布效应大小的多重检验,表明当样本量 $ n $ 至少以 $ \log m $ 的速率增长时,在温和正则性条件下,邦费罗尼校正与邦费罗尼-霍赫伯格程序(Benjamini-Hochberg procedure)在渐近意义上为贝叶斯最优。关键结果表明,当 $ n \propto \log m $ 时,邦费罗尼校正的 $ \alpha $ 固定,而 BH 程序则自适应于稀疏性,该结果基于具有未知效应大小分布的两组模型。
Recent results concerning asymptotic Bayes-optimality under sparsity (ABOS) of multiple testing procedures are extended to fairly generally distributed effect sizes under the alternative. An asymptotic framework is considered where both the number of tests m and the sample size m go to infinity, while the fraction p of true alternatives converges to zero. It is shown that under mild restrictions on the loss function nontrivial asymptotic inference is possible only if n increases to infinity at least at the rate of log m. Based on this assumption precise conditions are given under which the Bonferroni correction with nominal Family Wise Error Rate (FWER) level alpha and the Benjamini- Hochberg procedure (BH) at FDR level alpha are asymptotically optimal. When n is proportional to log m then alpha can remain fixed, whereas when n increases to infinity at a quicker rate, then alpha has to converge to zero roughly like n^(-1/2). Under these conditions the Bonferroni correction is ABOS in case of extreme sparsity, while BH adapts well to the unknown level of sparsity. In the second part of this article these optimality results are carried over to model selection in the context of multiple regression with orthogonal regressors. Several modifications of Bayesian Information Criterion are considered, controlling either FWER or FDR, and conditions are provided under which these selection criteria are ABOS. Finally the performance of these criteria is examined in a brief simulation study.
研究动机与目标
- 将稀疏性下的渐近贝叶斯最优性(ABOS)扩展至备择假设下效应大小为一般分布的多重检验程序,而不仅限于正态分布。
- 确定经典程序如邦费罗尼校正与邦费罗尼-霍赫伯格(BH)程序在高维稀疏设定下,且效应大小分布为一般分布时,实现 ABOS 的精确条件。
- 阐明样本量 $ n $ 与稀疏度水平 $ p \propto m^{-\beta} $ 在实现非平凡渐近推断与最优性中的作用。
- 将 ABOS 结果扩展至正交回归设计下多重回归中的模型选择,使用修正的 BIC 准则以控制家族错误率(FWER)或错误发现率(FDR)。
- 通过模拟研究验证理论发现,考察在不同稀疏度水平与样本量下,误分类率、FDR 与统计功效的表现。
提出的方法
- 采用渐近框架,其中检验数量 $ m $ 与样本量 $ n $ 同时趋于无穷,稀疏度 $ p \to 0 $,并假设 $ n \propto \log m $ 以实现非平凡推断。
- 应用两组模型,其中备择假设下的效应大小 $ \mu_i $ 从固定分布 $ \nu(\mu) $ 中抽取,允许一般(非正态)分布。
- 利用极值理论与高斯近似推导第一类与第二类错误概率的渐近行为,特别分析检验统计量的尾部行为。
- 通过证明在 $ \alpha $、$ p $ 与 $ \nu(\mu) $ 的特定条件下,某程序的贝叶斯风险与贝叶斯最优解之比收敛于 1,从而建立 ABOS。
- 将 ABOS 框架适配至多重回归,通过分析控制家族错误率(FWER)或错误发现率(FDR)的修正 BIC 准则,适用于正交设计。
- 采用模拟研究评估在不同稀疏度水平与样本量下的表现,指标包括误分类概率、FDR 与统计功效。
实验结果
研究问题
- RQ1当效应大小为一般分布且不一定是正态分布时,邦费罗尼校正在何种条件下为渐近贝叶斯最优?
- RQ2邦费罗尼-霍赫伯格程序在一般效应大小分布下是否能保持渐近贝叶斯最优性?其如何自适应于未知稀疏度水平?
- RQ3在稀疏高维检验中,样本量 $ n $ 相对于 $ m $ 的最小增长率是多少,才能实现非平凡渐近推断?
- RQ4当 $ n \propto \log m $ 与更快增长时,名义错误率 $ \alpha $ 的选择如何影响渐近最优性?
- RQ5在正交设计下,控制 FWER 或 FDR 的修正 BIC 准则在高维回归模型选择中是否为渐近贝叶斯最优?
主要发现
- 当 $ p \propto m^{-1} $ 且损失比 $ \delta \to 0 $ 的速率满足 $ \log \delta = o(\log m) $ 时,邦费罗尼校正为渐近贝叶斯最优(ABOS),且 $ \alpha $ 固定。
- 邦费罗尼-霍赫伯格程序在相同条件下亦为 ABOS,且对未知稀疏度水平具有良好的自适应性,即使 $ p \propto m^{-\beta} $ 且 $ \beta < 1 $ 时仍保持最优性。
- 仅当 $ n \to \infty $ 至少以 $ \log m $ 的速率增长时,才能实现非平凡渐近推断;在此条件下,若 $ n \propto \log m $,则 $ \alpha $ 可保持固定。
- 当 $ n $ 的增长快于 $ \log m $ 时,为使两类程序保持 ABOS,$ \alpha $ 必须大致以 $ n^{-1/2} $ 的速率收敛于零。
- 在正交设计下,用于高维回归模型选择的修正 BIC 准则,若控制 FWER 或 FDR,则在 $ n $、$ m $ 与 $ \alpha $ 的类似条件下为 ABOS,且对 $ \alpha $ 的依赖关系相似。
- 模拟结果证实,邦费罗尼校正与 BH 程序在不同稀疏度水平下均保持低误分类率与受控的 FDR,且在稀疏情形下 BH 展现出更优的统计功效。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。