[Paper Review] Comparing a Large Number of Multivariate Distributions
This paper proposes a nonparametric K-sample test for multivariate distributions using kernel mean embeddings, focusing on the maximum of pairwise maximum mean discrepancies (MMD) to detect sparse alternatives. The test is implemented via permutation resampling, with theoretical guarantees showing it converges to a Gumbel distribution under the null and achieves minimax rate optimality against sparse alternatives, outperforming average-type tests in high-dimensional, high-K settings.
In this paper, we propose a test for the equality of multiple distributions based on kernel mean embeddings. Our framework provides a flexible way to handle multivariate or even high-dimensional data by virtue of kernel methods and allows the number of distributions to increase with the sample size. This is in contrast to previous studies that have been mostly restricted to classical low-dimensional settings with a fixed number of distributions. By building on Cramer-type moderate deviation for degenerate two-sample V-statistics, we derive the limiting null distribution of the test statistic and show that it converges to a Gumbel distribution. The limiting distribution, however, depends on an infinite number of nuisance parameters, which makes it infeasible for use in practice. To address this issue, the proposed test is implemented via the permutation procedure and is shown to be minimax rate optimal against sparse alternatives. During our analysis, an exponential concentration inequality for the permuted test statistic is developed which may be of independent interest.
Motivation & Objective
- To address the lack of powerful multivariate K-sample tests for sparse alternatives where only a few distributions differ.
- To develop a test that remains effective as both sample size N and number of distributions K grow large.
- To overcome the intractability of the limiting null distribution, which depends on infinitely many nuisance parameters.
- To establish theoretical optimality and uniform consistency of the permutation-based test under sparse alternatives.
- To provide a practical, computationally feasible method for high-dimensional, high-K multivariate testing.
Proposed method
- The test statistic is defined as the maximum of pairwise MMDs between all K distributions, leveraging kernel mean embeddings for multivariate nonparametric comparison.
- The limiting null distribution of the test statistic is derived using Cramér-type moderate deviation for degenerate two-sample V-statistics, showing convergence to a Gumbel distribution.
- Due to the intractability of the limiting null distribution, the test uses a permutation procedure to determine critical values and p-values.
- An exponential concentration inequality for the permuted test statistic is developed using Bobkov’s inequality, relying only on observable quantities without moment assumptions.
- The permutation test is shown to be uniformly consistent and minimax rate optimal under regularity conditions for sparse alternatives.
- Simulations use randomized permutations with M=200 draws and estimate power over 800 repetitions at α=0.05.
Experimental results
Research questions
- RQ1Can a multivariate K-sample test be developed that maintains high power when only a small number of distributions differ (sparse alternatives), even as K increases?
- RQ2What is the limiting null distribution of a maximum-type test statistic based on kernel MMD in a high-K, high-N asymptotic regime?
- RQ3How can the intractable infinite-dimensional nuisance parameter problem in the limiting null distribution be overcome in practice?
- RQ4Is the permutation-based implementation of the test uniformly consistent and minimax optimal against sparse alternatives?
- RQ5How does the proposed maximum-type test compare in power to average-type MMD tests in high-dimensional, sparse settings?
Key findings
- The proposed test based on the maximum pairwise MMD achieves power close to one even when K reaches 100, demonstrating robustness to increasing numbers of distributions in sparse scenarios.
- In contrast, average-type tests like DISCO and ECF show decreasing power as K increases, due to dilution of the signal in the average pairwise distance.
- The limiting null distribution of the test statistic converges to a Gumbel distribution under the null hypothesis, though it depends on infinitely many nuisance parameters.
- The permutation procedure is uniformly consistent and achieves minimax rate optimality against sparse alternatives under regularity conditions.
- An exponential concentration inequality for the permuted test statistic is derived using Bobkov’s inequality, providing a sharp tail bound without moment assumptions.
- The method is shown to be particularly effective in high-dimensional settings (d=5) with fixed sample sizes per group (n=10), outperforming existing methods in sparse signal detection.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.