[Paper Review] A Kernel Two-Sample Test for Functional Data
Proposes a nonparametric kernel-based two-sample test for comparing distributions of functional data using MMD on function spaces, with theory for kernels on Hilbert spaces and scaling analysis for discretised observations.
We propose a nonparametric two-sample test procedure based on Maximum Mean Discrepancy (MMD) for testing the hypothesis that two samples of functions have the same underlying distribution, using kernels defined on function spaces. This construction is motivated by a scaling analysis of the efficiency of MMD-based tests for datasets of increasing dimension. Theoretical properties of kernels on function spaces and their associated MMD are established and employed to ascertain the efficacy of the newly proposed test, as well as to assess the effects of using functional reconstructions based on discretised function samples. The theoretical results are demonstrated over a range of synthetic and real world datasets.
Motivation & Objective
- Motivate nonparametric two-sample testing for functional data arising from discretised functions.
- Generalise kernel-based MMD testing to real, separable Hilbert spaces to handle functional data.
- Establish conditions for kernels to be characteristic on function spaces and describe the associated RKHS.
- Analyze how discretisation (mesh size) affects test power and how to mitigate it through kernel scaling.
- Demonstrate theoretical properties and empirical performance on synthetic and real datasets.
Proposed method
- Define and study kernels on real, separable Hilbert spaces and their RKHS.
- Introduce and use Maximum Mean Discrepancy (MMD) as a statistic for two-sample testing on function spaces.
- Provide closed-form MMD expressions and unbiased estimators (U-statistic and linear-time variant).
- Analyze scaling of kernel bandwidth with mesh size to achieve power that is independent of discretisation for Gaussian processes.
- Construct and characterise a squared-exponential kernel on function spaces (SE-T) and derive RKHS descriptions.
- Discuss implications of using reconstructed functional data and establish connections to weak convergence.
Experimental results
Research questions
- RQ1What conditions make kernels on function spaces characteristic, ensuring MMD is a metric for distributions over functions?
- RQ2How does discretisation (mesh size) influence the power of kernel two-sample tests for functional data, and can kernel scaling mitigate this?
- RQ3How can kernels be defined directly on Hilbert spaces of functions, and what is the structure of their RKHS?
- RQ4What are the asymptotic distributions and power properties of MMD estimators (unbiased and linear-time) in the functional data setting?
- RQ5How do Gaussian process assumptions help derive closed-form MMD expressions and scaling laws for testing?
Key findings
- Kernel two-sample tests based on MMD can be constructed on function spaces with characteristic kernels ensuring valid testing.
- Under mean-shift alternatives, the MMD-based test power can be made asymptotically mesh-size independent via appropriate bandwidth scaling.
- A broad class of kernels on real separable Hilbert spaces is developed, with explicit RKHS characterization for a squared-exponential type kernel on Hilbert spaces (SE-T).
- Reconstruction of discretised functional data affects the test, and theoretical results quantify these impacts.
- The paper provides both theoretical results and numerical experiments comparing the kernel-based test to existing functional data two-sample tests, validating the scaling and efficacy.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.