Skip to main content
QUICK REVIEW

[Paper Review] Multiscale Fisher's Independence Test for Multivariate Dependence

Shai Gorsky, Li Ma|arXiv (Cornell University)|Jun 18, 2018
Statistical Methods and Inference20 references4 citations
TL;DR

This paper proposes Multiscale Fisher's Independence Test (MultiFIT), a scalable, resampling-free method for testing multivariate dependence by decomposing the problem into sequential univariate independence tests on $2\times2$ contingency tables through coarse-to-fine discretization. It achieves finite-sample level control and strong consistency with near-linear computational complexity, enabling efficient inference on massive datasets while learning underlying dependency structures.

ABSTRACT

Identifying dependency in multivariate data is a common inference task that arises in numerous applications. However, existing nonparametric independence tests typically require computation that scales at least quadratically with the sample size, making it difficult to apply them to massive data. Moreover, resampling is usually necessary to evaluate the statistical significance of the resulting test statistics at finite sample sizes, further worsening the computational burden. We introduce a scalable, resampling-free approach to testing the independence between two random vectors by breaking down the task into simple univariate tests of independence on a collection of 2x2 contingency tables constructed through sequential coarse-to-fine discretization of the sample space, transforming the inference task into a multiple testing problem that can be completed with almost linear complexity with respect to the sample size. To address increasing dimensionality, we introduce a coarse-to-fine sequential adaptive procedure that exploits the spatial features of dependency structures to more effectively examine the sample space. We derive a finite-sample theory that guarantees the inferential validity of our adaptive procedure at any given sample size. In particular, we show that our approach can achieve strong control of the family-wise error rate without resampling or large-sample approximation. We demonstrate the substantial computational advantage of the procedure in comparison to existing approaches as well as its decent statistical power under various dependency scenarios through an extensive simulation study, and illustrate how the divide-and-conquer nature of the procedure can be exploited to not just test independence but to learn the nature of the underlying dependency. Finally, we demonstrate the use of our method through analyzing a large data set from a flow cytometry experiment.

Motivation & Objective

  • To address the computational infeasibility of existing nonparametric independence tests on massive multivariate datasets.
  • To eliminate reliance on resampling or asymptotic approximations for significance testing in finite samples.
  • To develop a method that maintains exact level control while scaling efficiently with sample size.
  • To exploit spatial structure in dependency to reduce the number of required univariate tests.
  • To enable not only independence testing but also learning of the nature of multivariate dependencies.

Proposed method

  • The method transforms multivariate independence testing into a multiple testing problem over $2\times2$ contingency tables formed by sequential coarse-to-fine discretization of the sample space.
  • It employs a coarse-to-fine sequential adaptive procedure that selectively tests only relevant scales, reducing computational burden.
  • Fisher’s exact test is applied to each $2\times2$ table, with p-values adjusted using the mid-p correction for improved inference.
  • The procedure uses a closed testing approach to control the family-wise error rate and ensures finite-sample validity.
  • The algorithm dynamically selects resolution levels based on data-adaptive criteria, focusing on regions with potential dependence.
  • The maximal resolution is set to $\lfloor \log_2(n/10) \rfloor$, where $n$ is the sample size.

Experimental results

Research questions

  • RQ1Can a nonparametric multivariate independence test be designed with near-linear computational complexity?
  • RQ2Can finite-sample level control be achieved without resampling or asymptotic approximation?
  • RQ3Can data-adaptive, coarse-to-fine discretization reduce the number of univariate tests while preserving power?
  • RQ4Does the method retain strong consistency in large samples?
  • RQ5Can the divide-and-conquer structure reveal the spatial nature of multivariate dependencies?

Key findings

  • MultiFIT achieves finite-sample level control at any given sample size without resampling or asymptotic approximation, as confirmed by simulations with up to 2000 observations.
  • The method demonstrates substantial computational speedup compared to existing approaches, with runtimes scaling nearly linearly with sample size across all scenarios.
  • Under various dependency structures—including linear, parabolic, and local dependencies—MultiFIT maintains robust statistical power, especially when tuned with $p^* \geq 0.05$ and $R^* \geq 2$.
  • The adaptive procedure reduces the number of tests significantly compared to exhaustive search, particularly in high-dimensional settings.
  • In the flow cytometry application, MultiFIT successfully identified biologically relevant dependencies missed by standard methods.
  • The method's ability to localize dependency structures was validated in simulation scenarios with embedded signals, where it outperformed alternatives in detecting local dependencies.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.