[Paper Review] A Kernel Test of Goodness of Fit
We propose a nonparametric goodness-of-fit test using a Stein discrepancy in an RKHS and compute its null distribution via a wild bootstrap, applicable to i.i.d. and dependent samples.
We propose a nonparametric statistical test for goodness-of-fit: given a set of samples, the test determines how likely it is that these were generated from a target density function. The measure of goodness-of-fit is a divergence constructed via Stein's method using functions from a Reproducing Kernel Hilbert Space. Our test statistic is based on an empirical estimate of this divergence, taking the form of a V-statistic in terms of the log gradients of the target density and the kernel. We derive a statistical test, both for i.i.d. and non-i.i.d. samples, where we estimate the null distribution quantiles using a wild bootstrap procedure. We apply our test to quantifying convergence of approximate Markov Chain Monte Carlo methods, statistical model criticism, and evaluating quality of fit vs model complexity in nonparametric density estimation.
Motivation & Objective
- Develop a nonparametric goodness-of-fit test based on Stein’s method within an RKHS framework.
- Avoid reliance on target density integrals by using kernels and gradients of log-target density.
- Provide a practical statistical test with bootstrap-calibrated thresholds for both independent and dependent samples.
- Demonstrate applications to approximate MCMC convergence, model criticism, and nonparametric density estimation.
Proposed method
- Define a Stein operator in RKHS and derive a closed-form Stein discrepancy S_p(Z) as the RKHS norm of E_q[ξ_p(Z)].
- Express the discrepancy via a symmetric kernel function h_p and show S_p^2(Z)=E_q[h_p(Z,Z')] with Z' independent of Z.
- Construct a quadratic-time V-statistic estimator V_n for S_p^2(Z) from samples {Z_i}.
- Use wild bootstrap to estimate null distribution quantiles for dependent data and derive a practical testing procedure.
- Prove that the kernel choice is universal under mild conditions, ensuring discrimination between p and q.
- Provide asymptotic results under tau-mixing for the null distribution and bootstrap validity.
Experimental results
Research questions
- RQ1Can a kernel-based Stein discrepancy distinguish any difference between the target distribution p and the observed distribution q?
- RQ2How can we reliably estimate the null distribution of the Stein-based test statistic for i.i.d. and dependent samples?
- RQ3Does the proposed test work for assessing approximate MCMC convergence, model criticism, and nonparametric density estimation?
- RQ4What are the practical guidelines for bootstrap tuning to handle correlation in data?
Key findings
- The test statistic S_p(Z) is given by the RKHS norm of E_q[ξ_p(Z)], with a closed-form h_p representation.
- Under certain conditions, S_p^2(Z)=E_q[h_p(Z,Z')] and the statistic distinguishes p from q when the kernel is C_0-universal.
- A wild bootstrap procedure provides consistently calibrated p-values for both independent and dependent samples.
- The method yields practical insights into approximate MCMC bias-variance, GP model criticism, and convergence of nonparametric density estimators.
- The approach does not require sampling from the target distribution or its normalization constant.
- Code for replication is available at the authors' repository.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.