Skip to main content
QUICK REVIEW

[Paper Review] The art of BART: Minimax optimality over nonhomogeneous smoothness in high dimension

Seonghyun Jeong, Veronika Ročková|arXiv (Cornell University)|Aug 15, 2020
Statistical Methods and Inference4 citations
TL;DR

This paper establishes minimax optimality of Bayesian additive regression trees (BART) under nonhomogeneous, anisotropic smoothness and discontinuities in high-dimensional function estimation. By introducing a novel class of sparse piecewise heterogeneous anisotropic Hölder functions and leveraging Dirichlet subset selection priors, the authors prove that BART achieves optimal posterior contraction rates without requiring isotropy or homogeneity assumptions, outperforming Gaussian processes and other default ML tools in complex, real-world scenarios.

ABSTRACT

Many asymptotically minimax procedures for function estimation often rely on somewhat arbitrary and restrictive assumptions such as isotropy or spatial homogeneity. This work enhances the theoretical understanding of Bayesian additive regression trees under substantially relaxed smoothness assumptions. We provide a comprehensive study of asymptotic optimality and posterior contraction of Bayesian forests when the regression function has anisotropic smoothness that possibly varies over the function domain. The regression function can also be possibly discontinuous. We introduce a new class of sparse {\em piecewise heterogeneous anisotropic} Hölder functions and derive their minimax lower bound of estimation in high-dimensional scenarios under the $L_2$-loss. We then find that the Bayesian tree priors, coupled with a Dirichlet subset selection prior for sparse estimation in high-dimensional scenarios, adapt to unknown heterogeneous smoothness, discontinuity, and sparsity. These results show that Bayesian forests are uniquely suited for more general estimation problems that would render other default machine learning tools, such as Gaussian processes, suboptimal. Our numerical study shows that Bayesian forests often outperform other competitors such as random forests and deep neural networks, which are believed to work well for discontinuous or complicated smooth functions. Beyond nonparametric regression, we also examined posterior contraction of Bayesian forests for density estimation and binary classification using the technique developed in this study.

Motivation & Objective

  • To close the theoretical gap in Bayesian forests by analyzing their performance under relaxed smoothness assumptions, particularly anisotropic and inhomogeneous smoothness.
  • To address the disconnect between existing asymptotic minimax theory—often based on restrictive isotropy or homogeneity—and real-world data, which frequently exhibit spatially varying smoothness.
  • To develop a new function class that captures piecewise, heterogeneous, anisotropic smoothness with possible discontinuities, enabling more realistic nonparametric function estimation.
  • To prove that BART with Dirichlet subset selection priors achieves posterior contraction at minimax-optimal rates across this general function class.
  • To extend theoretical results beyond regression to density estimation and binary classification, demonstrating broad applicability of the framework.

Proposed method

  • Introduce a new class of functions—piecewise heterogeneous anisotropic Hölder functions—where each rectangular subdomain has its own anisotropic smoothness level, with smoothness possibly varying across regions and allowing discontinuities.
  • Derive a minimax lower bound for estimation under $L_2$-loss in high-dimensional settings for this new function class, establishing the theoretical benchmark for optimality.
  • Use a Dirichlet prior on tree partitioning and a sparse subset selection prior on regression coefficients to induce sparsity and adaptivity to unknown smoothness and discontinuities.
  • Apply posterior contraction theory to show that the BART posterior concentrates at the minimax rate, even when smoothness varies spatially and is unknown a priori.
  • Leverage stick-breaking representations of Dirichlet processes to bound the probability of sparse recovery, ensuring that the prior favors low-dimensional, relevant features.
  • Establish theoretical bounds on posterior concentration using concentration inequalities and properties of the Gamma function, particularly for small concentration parameters in high dimensions.

Experimental results

Research questions

  • RQ1Can Bayesian forests achieve minimax optimal estimation rates under nonhomogeneous, anisotropic smoothness in high-dimensional regression?
  • RQ2How does BART perform when the regression function exhibits discontinuities or varying smoothness across different regions of the input space?
  • RQ3Can the posterior contraction rate of BART be proven to match the minimax lower bound for a general class of piecewise heterogeneous anisotropic functions?
  • RQ4Does the combination of tree-based priors and Dirichlet subset selection enable adaptive estimation of sparsity, discontinuities, and heterogeneous smoothness without prior knowledge?
  • RQ5To what extent do the theoretical results extend beyond regression to density estimation and binary classification?

Key findings

  • The paper establishes a minimax lower bound for $L_2$-risk in high-dimensional nonparametric regression over a new class of piecewise heterogeneous anisotropic Hölder functions with possible discontinuities.
  • BART with Dirichlet subset selection priors achieves posterior contraction at the minimax-optimal rate, even when smoothness varies across regions and directions.
  • The posterior probability of being within $\epsilon$-neighborhood of the true function $\eta^*$ is bounded below by $\exp\{-C\xi s\log(p/\epsilon)\}$, confirming effective concentration.
  • The prior ensures sparsity by assigning exponentially small probability to configurations where more than $s$ variables are active, with $\Pi(\min_{|S|=s} \sum_{j\notin S} \eta_j \geq \epsilon) \leq \exp\{-C(\xi-1)s\log p - \log \epsilon\}$.
  • Theoretical results show that BART adapts to unknown smoothness, discontinuities, and sparsity without tuning, outperforming Gaussian processes and other default ML tools in complex settings.
  • Numerical studies confirm that BART outperforms random forests and deep neural networks in estimating discontinuous or highly heterogeneous smooth functions, validating theoretical claims empirically.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.