Skip to main content
QUICK REVIEW

[论文解读] Multiscale Fisher's Independence Test for Multivariate Dependence

Shai Gorsky, Li Ma|arXiv (Cornell University)|Jun 18, 2018
Statistical Methods and Inference参考文献 20被引用 4
一句话总结

本文提出多尺度费雪独立性检验(MultiFIT),一种可扩展、无需重采样的方法,通过从粗到细的离散化将多变量依赖性检验问题分解为一系列顺序的单变量独立性检验,基于 $2\times2$ 列联表实现。该方法在有限样本下实现显著性水平控制与强一致性,计算复杂度接近线性,可在大规模数据集上高效进行推断,同时学习潜在的依赖结构。

ABSTRACT

Identifying dependency in multivariate data is a common inference task that arises in numerous applications. However, existing nonparametric independence tests typically require computation that scales at least quadratically with the sample size, making it difficult to apply them to massive data. Moreover, resampling is usually necessary to evaluate the statistical significance of the resulting test statistics at finite sample sizes, further worsening the computational burden. We introduce a scalable, resampling-free approach to testing the independence between two random vectors by breaking down the task into simple univariate tests of independence on a collection of 2x2 contingency tables constructed through sequential coarse-to-fine discretization of the sample space, transforming the inference task into a multiple testing problem that can be completed with almost linear complexity with respect to the sample size. To address increasing dimensionality, we introduce a coarse-to-fine sequential adaptive procedure that exploits the spatial features of dependency structures to more effectively examine the sample space. We derive a finite-sample theory that guarantees the inferential validity of our adaptive procedure at any given sample size. In particular, we show that our approach can achieve strong control of the family-wise error rate without resampling or large-sample approximation. We demonstrate the substantial computational advantage of the procedure in comparison to existing approaches as well as its decent statistical power under various dependency scenarios through an extensive simulation study, and illustrate how the divide-and-conquer nature of the procedure can be exploited to not just test independence but to learn the nature of the underlying dependency. Finally, we demonstrate the use of our method through analyzing a large data set from a flow cytometry experiment.

研究动机与目标

  • 解决现有非参数独立性检验在大规模多变量数据集上计算不可行的问题。
  • 消除在有限样本中进行显著性检验时对重采样或渐近近似的依赖。
  • 开发一种在样本量增大时仍能保持精确显著性水平控制且计算效率高的方法。
  • 利用依赖关系中的空间结构,减少所需单变量检验的次数。
  • 不仅实现独立性检验,还能揭示多变量依赖关系的本质。

提出的方法

  • 该方法将多变量独立性检验转化为在通过样本空间顺序从粗到细离散化生成的 $2\times2$ 列联表上的多重检验问题。
  • 采用从粗到细的顺序自适应过程,仅选择性地测试相关尺度,从而减轻计算负担。
  • 对每个 $2\times2$ 表应用费雪精确检验,并使用中位p值校正以提升推断性能。
  • 通过封闭检验方法控制家庭错误率,确保有限样本下的有效性。
  • 该算法基于数据自适应标准动态选择分辨率层级,聚焦于可能存在依赖的区域。
  • 最大分辨率设定为 $\lfloor \log_2(n/10) \rfloor$,其中 $n$ 为样本量。

实验结果

研究问题

  • RQ1能否设计一种具有近似线性计算复杂度的非参数多变量独立性检验?
  • RQ2能否在不依赖重采样或渐近近似的情况下实现在有限样本下的显著性水平控制?
  • RQ3能否通过数据自适应的从粗到细离散化方法,减少单变量检验次数,同时保持检验效能?
  • RQ4该方法在大样本下是否仍保持强一致性?
  • RQ5其分治结构能否揭示多变量依赖关系的空间特性?

主要发现

  • 通过模拟验证,MultiFIT 在任意给定样本量下均能实现有限样本下的显著性水平控制,无需重采样或渐近近似,模拟样本量最高达 2000 个。
  • 与现有方法相比,该方法在所有场景下均展现出显著的计算加速,运行时间几乎随样本量线性增长。
  • 在多种依赖结构下——包括线性、抛物线型及局部依赖——MultiFIT 保持了稳健的统计功效,尤其在 $p^* \geq 0.05$ 且 $R^* \geq 2$ 的调参条件下表现更优。
  • 自适应过程显著减少了所需检验次数,尤其在高维设置中优势明显。
  • 在流式细胞术应用中,MultiFIT 有效识别出标准方法遗漏的生物相关依赖关系。
  • 在嵌入信号的模拟场景中,该方法对依赖结构的定位能力得到验证,其在检测局部依赖关系方面优于其他对比方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。