Skip to main content
QUICK REVIEW

[Paper Review] Nonparametric Bayesian Aggregation for Massive Data

Zuofeng Shang, Botao Hao|arXiv (Cornell University)|Aug 17, 2015
Statistical Methods and Inference30 references3 citations
TL;DR

This paper proposes a scalable nonparametric Bayesian inference framework for massive datasets by splitting data into subsets, performing independent Gaussian process posterior inference on each, and analytically aggregating local results into global estimates without additional computation. The method achieves oracle-like performance: aggregated credible balls match the radius, credibility, and frequentist coverage of those computed from the full data, under proper prior and subset size choices.

ABSTRACT

We develop a set of scalable Bayesian inference procedures for a general class of nonparametric regression models. Specifically, nonparametric Bayesian inferences are separately performed on each subset randomly split from a massive dataset, and then the obtained local results are aggregated into global counterparts. This aggregation step is explicit without involving any additional computation cost. By a careful partition, we show that our aggregated inference results obtain an oracle rule in the sense that they are equivalent to those obtained directly from the entire data (which are computationally prohibitive). For example, an aggregated credible ball achieves desirable credibility level and also frequentist coverage while possessing the same radius as the oracle ball.

Motivation & Objective

  • To develop a computationally efficient Bayesian inference procedure for massive nonparametric regression data.
  • To address uncertainty quantification in nonparametric Bayesian models under data size constraints.
  • To ensure aggregated posterior results maintain both Bayesian and frequentist validity (e.g., coverage and credibility).
  • To establish theoretical guarantees for aggregation under a uniform Bernstein-von Mises theorem.
  • To minimize computational cost while preserving statistical optimality in estimation and credible region construction.

Proposed method

  • The dataset is randomly partitioned into multiple subsets to enable distributed computation.
  • Nonparametric Bayesian regression is performed independently on each subset using a Gaussian process prior.
  • Posterior means and credible balls are computed via MCMC on each subset, approximating functionals through finite Fourier expansions.
  • Global posterior mean is formed by weighted averaging of individual posterior mode Fourier coefficients.
  • Global credible ball radius is computed via an explicit formula combining individual radii, avoiding re-computation.
  • The aggregation is fully analytical, incurring no additional computational cost beyond local inference.

Experimental results

Research questions

  • RQ1Can distributed Bayesian inference for massive nonparametric regression achieve the same statistical performance as inference on the full dataset?
  • RQ2How can local posterior distributions be aggregated to yield a global posterior with optimal frequentist coverage and credible level?
  • RQ3What conditions on the number of subsets and prior choice ensure the aggregated credible ball matches the oracle ball in radius and coverage?
  • RQ4Can the aggregation process preserve the optimal estimation rate of the full-data posterior mean?
  • RQ5What theoretical foundation supports the uniform convergence of local posteriors to Gaussian processes across subsets?

Key findings

  • The aggregated posterior mean achieves the optimal nonparametric estimation rate, matching the oracle rate from the full dataset.
  • The aggregated credible ball attains the same asymptotic radius as the oracle credible ball under proper prior and subset size choices.
  • The aggregated credible ball maintains both the desired credibility level (1−α) and frequentist coverage probability (1−α) under the true model.
  • The method achieves an oracle rule in the sense that aggregated results are statistically equivalent to those obtained from the entire dataset, despite being computed distributively.
  • The theoretical foundation relies on a uniform nonparametric Gaussian approximation theorem (uniform Bernstein-von Mises), ensuring local posteriors converge to Gaussian processes uniformly over the number of subsets.
  • Computational speed is dramatically improved with increasing data size, as aggregation is analytical and incurs no additional cost beyond local MCMC sampling.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.