Skip to main content
QUICK REVIEW

[Paper Review] Measuring Association on Topological Spaces Using Kernels and Geometric Graphs

Nabarun Deb, Promit Ghosal|arXiv (Cornell University)|Oct 5, 2020
Statistical Methods and Inference99 references31 citations
TL;DR

Nonparametric measures of association for X and Y in general topological spaces are built from RKHS kernels and geometric graphs, enabling independence testing and adaptation to intrinsic dimensionality.

ABSTRACT

In this paper we propose and study a class of simple, nonparametric, yet interpretable measures of association between two random variables $X$ and $Y$ taking values in general topological spaces. These nonparametric measures -- defined using the theory of reproducing kernel Hilbert spaces -- capture the strength of dependence between $X$ and $Y$ and have the property that they are 0 if and only if the variables are independent and 1 if and only if one variable is a measurable function of the other. Further, these population measures can be consistently estimated using the general framework of graph functionals which include $k$-nearest neighbor graphs and minimum spanning trees. Moreover, a sub-class of these estimators are also shown to adapt to the intrinsic dimensionality of the underlying distribution. Some of these empirical measures can also be computed in near linear time. Under the hypothesis of independence between $X$ and $Y$, these empirical measures (properly normalized) have a standard normal limiting distribution. Thus, these measures can also be readily used to test the hypothesis of mutual independence between $X$ and $Y$. In fact, as far as we are aware, these are the only procedures that possess all the above mentioned desirable properties. Furthermore, when restricting to Euclidean spaces, we can make these sample measures of association finite-sample distribution-free, under the hypothesis of independence, by using multivariate ranks defined via the theory of optimal transport. The recent correlation coefficient proposed in Dette et al. (2013), Chatterjee (2019), and Azadkia and Chatterjee (2019) can be seen as a special case of this general class of measures.

Motivation & Objective

  • Define population and empirical measures of association for (X,Y) in general topological spaces.
  • Develop a kernel-based measure that equals 0 iff X and Y are independent and 1 iff Y is a noiseless function of X.
  • Provide consistent estimators (KMAc) using geometric graphs like k-NN graphs and MSTs.
  • Establish limiting normality under independence and derive time-efficient estimators.
  • Show adaptation to intrinsic dimensionality and discuss finite-sample properties in Euclidean settings.

Proposed method

  • Introduce eta_K as a population measure using a characteristic kernel K on Y and RKHS H_K.
  • Construct empirical estimators etâ_n via geometric graphs G_n on X with edges linking nearby X_i, using K(Y_i,Y_j).
  • Define the kernel measure of association estimator η̂_n = [ (1/n) sum_i d_i^{-1} sum_{j:(i,j)∈E(G_n)} K(Y_i,Y_j) - (1/[n(n-1)]) sum_{i≠j} K(Y_i,Y_j) ] / [ (1/n) sum_i K(Y_i,Y_i) - (1/[n(n-1)]) sum_{i≠j} K(Y_i,Y_j) ].
  • Explain conditions (A1)-(A2) on G_n for consistency to η_K(μ), and show η_K(μ) = 1 - E||K(·,Y′)−K(·,Ỹ′)||_H^2 / E||K(·,Y)−K(·,Y′)||_H^2.
  • Discuss a linear-time variant η̂_n^{lin} and a rank-based, distribution-free version η̂_n^{rank} using multivariate ranks from optimal transport.
  • Address computational aspects with near-linear time implementations and CLTs under independence.

Experimental results

Research questions

  • RQ1Can one define a simple, nonparametric measure of association for X and Y that is 0 under independence and 1 when Y is a function of X in general topological spaces?
  • RQ2How can RKHS kernels and geometric graphs be combined to estimate this association consistently from data?
  • RQ3What are the asymptotic distributional properties (e.g., CLT) of the proposed measures under independence, and how do they enable testing?
  • RQ4Can the proposed methods adapt to intrinsic dimensionality and achieve near-linear computational complexity?

Key findings

  • A population kernel measure η_K(μ) is defined that satisfies 0 for independence and 1 for noiseless functional dependence, under a characteristic kernel K.
  • Empirical estimators η̂_n based on k-NN graphs and other geometric graphs consistently estimate η_K(μ).
  • Under independence, η̂_n (properly normalized) satisfies a standard normal CLT uniformly over a large class of graphs.
  • The method adapts to intrinsic dimensionality of X and Y, and a near-linear time estimator η̂_n^{lin} is available.
  • When X and Y are Euclidean, a distribution-free variant η̂_n^{rank} using multivariate ranks provides finite-sample distribution-free testing under independence.
  • The framework includes and generalizes correlation-type measures such as those proposed by Dette et al., Chatterjee, and Azadkia–Chatterjee as special cases.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.