[论文解读] Selective inference for the problem of regions via multiscale bootstrap
本文提出了一种用于多元正态模型中假设检验的筛选推断方法,聚焦于参数空间中的任意形状区域——特别是层次聚类问题。通过利用多尺度自展法重采样和基于有符号距离与平均曲率的几何校正,该方法生成了渐近第二阶准确的p值,有效校正了选择偏差,且计算效率与迭代自展法相当,但成本更低。
A general approach to selective inference is considered for hypothesis testing of the null hypothesis represented as an arbitrary shaped region in the parameter space of multivariate normal model. This approach is useful for hierarchical clustering where confidence levels of clusters are calculated only for those appeared in the dendrogram, thus subject to heavy selection bias. Our computation is based on a raw confidence measure, called bootstrap probability, which is easily obtained by counting how many times the same cluster appears in bootstrap replicates of the dendrogram. We adjust the bias of the bootstrap probability by utilizing the scaling-law in terms of geometric quantities of the region in the abstract parameter space, namely, signed distance and mean curvature. Although this idea has been used for non-selective inference of hierarchical clustering, its selective inference version has not been discussed in the literature. Our bias-corrected $p$-values are asymptotically second-order accurate in the large sample theory of smooth boundary surfaces of regions, and they are also justified for nonsmooth surfaces such as polyhedral cones. The $p$-values are asymptotically equivalent to those of the iterated bootstrap but with less computation.
研究动机与目标
- 为解决通过层次聚类识别的聚类在频率学推断中因数据依赖的聚类选择而产生的严重选择偏差这一关键问题。
- 为多元正态参数空间中任意形状区域的筛选推断建立一个通用框架。
- 将先前仅用于非筛选推断的多尺度自展法扩展至筛选推断场景。
- 提供一种计算效率更高的替代方案,与迭代自展法具有相似的渐近准确性,但成本更低。
- 确保在高维参数空间中对光滑与非光滑区域(如多面锥)均有效。
提出的方法
- 利用多尺度自展重采样估计参数空间中区域的几何量——有符号距离与平均曲率。
- 通过几何量的标度律对自展概率(BP)进行偏差校正,从而得到筛选推断的校正p值。
- 推导出一个p值 $ p_{ ext{SI},k}(H|S,y) $,在边界光滑条件下具有渐近第二阶准确性。
- 通过涉及不完全伽马函数的积分及归一化常数表示p值。
- 理论证明其与迭代自展p值等价,但计算成本更低。
- 在多元正态性假设及边界函数的Lipschitz连续性条件下验证方法的有效性。
实验结果
研究问题
- RQ1如何在层次聚类中一致地应用筛选推断,以应对因数据依赖的聚类选择而产生的严重选择偏差?
- RQ2多尺度自展法能否被调整以在筛选推断中提供有效的p值,而不仅限于非筛选推断?
- RQ3几何量——有符号距离与平均曲率——在筛选推断中校正自展概率偏差的过程中起什么作用?
- RQ4在筛选推断场景下,所提出方法与迭代自展法相比,在准确性和效率方面表现如何?
- RQ5该方法对非光滑边界(如聚类问题中常见的多面锥)是否具有鲁棒性?
主要发现
- 在多元正态模型中,边界光滑条件下,所提出的筛选推断p值具有渐近第二阶准确性。
- 该方法在计算效率上与标准自展法相当,显著低于迭代自展法的成本,同时保持了相似的渐近准确性。
- 偏差校正后的p值在理论上等价于迭代自展法的结果,通过不完全伽马函数的积分表示得以证明。
- 该方法对光滑与非光滑区域(包括多面锥)均有效,扩展了其在真实聚类问题中的适用性。
- 通过将p值表示为涉及 $ I_k( heta) $ 和 $ J_k( heta) $ 的积分,提供了理论依据,满足第二阶准确性的条件。
- 该方法有效校正了标准自展概率(如 pvclust)中观察到的严重选择偏差,如在肺部数据集示例中 $ p_{ ext{AU}} $ 与 $ p_{ ext{SI}} $ 之间存在显著差异所示。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。