[Paper Review] Large sample analysis of the median heuristic
The paper proves that the bandwidth chosen by the median heuristic is asymptotically normal in kernel two-sample testing, by deriving a CLT for a class of U-statistics and showing convergence of the empirical distance distribution to a target mixture. It also compares median-based bandwidth to power-maximization empirically.
In kernel methods, the median heuristic has been widely used as a way of setting the bandwidth of RBF kernels. While its empirical performances make it a safe choice under many circumstances, there is little theoretical understanding of why this is the case. Our aim in this paper is to advance our understanding of the median heuristic by focusing on the setting of kernel two-sample test. We collect new findings that may be of interest for both theoreticians and practitioners. In theory, we provide a convergence analysis that shows the asymptotic normality of the bandwidth chosen by the median heuristic in the setting of kernel two-sample test. Systematic empirical investigations are also conducted in simple settings, comparing the performances based on the bandwidths chosen by the median heuristic and those by the maximization of test power.
Motivation & Objective
- Understand why the median heuristic is effective in kernel methods, especially kernel two-sample tests.
- Characterize the asymptotic behavior of the median-based bandwidth as sample size grows.
- Develop a central limit theorem for a class of U-statistics underlying pairwise distances.
- Relate the empirical median of pairwise distances to the theoretical median of a target mixture distribution.
- Provide empirical insights comparing median-based bandwidth to bandwidths chosen by test-power maximization.
Proposed method
- Define the median heuristic as ν = sqrt(H_n/2) where H_n is the empirical median of pairwise squared distances.
- Model the pairwise squared distances as a three-component mixture corresponding to intra-X, intra-Y, and inter-XY distances with weights α^2, (1−α)^2, and 2α(1−α).
- Show that the empirical cdf of pairwise distances converges to a target cdf F and establish a CLT for the resulting U-statistic U_n.
- Prove the asymptotic normality of the squared sample median H_n via a CLT for non-identically distributed triangular-array U-statistics.
- Provide a formula for the asymptotic variance and connect it to the derivative F'(m) at the target median m.
- Conduct empirical comparisons between median-based bandwidth and power-maximization bandwidth in Gaussian-kernel two-sample tests.
Experimental results
Research questions
- RQ1Does the median heuristic converge to the median of a target distribution as sample size grows?
- RQ2Is the median-based bandwidth asymptotically normal, and if so, what is its asymptotic variance?
- RQ3How does the median heuristic perform relative to bandwidths chosen by maximization of test power in kernel two-sample tests?
- RQ4What is the effect of intra- and inter-segment distances on the distribution of pairwise distances and the median heuristic?
- RQ5Can the theory be extended to multiple change-points or asymmetric sample splits (α depending on n)?
Key findings
- The empirical distribution of pairwise squared distances converges to a target mixture distribution reflecting intra-X, intra-Y, and inter-XY distances.
- The median of the target distribution exists and, under a nonzero derivative condition, the empirical median H_n is asymptotically normal with a computable variance.
- A central limit theorem is established for a broad class of non-identically distributed U-statistics, underpinning the median empirical distribution result.
- The median-based bandwidth can closely match bandwidths chosen by power-maximization in certain scenarios, particularly mean shifts, but may be less optimal in others like variance changes.
- The paper provides an approximate Bahadur slope analysis for different test statistics and bandwidth choices, linking bandwidth to test efficiency.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.