Skip to main content
QUICK REVIEW

[论文解读] The art of BART: Minimax optimality over nonhomogeneous smoothness in high dimension

Seonghyun Jeong, Veronika Ročková|arXiv (Cornell University)|Aug 15, 2020
Statistical Methods and Inference被引用 4
一句话总结

该论文在高维函数估计中,针对非齐次、各向异性的光滑性及不连续性,建立了贝叶斯加性回归树(BART)的极小极大最优性。通过引入一类新型的稀疏分段异质各向异性霍尔德函数,并利用狄利克雷子集选择先验,作者证明了BART在无需假设各向同性或齐次性的情况下,可实现最优后验收缩速率,优于高维复杂现实场景中的高斯过程及其他默认机器学习工具。

ABSTRACT

Many asymptotically minimax procedures for function estimation often rely on somewhat arbitrary and restrictive assumptions such as isotropy or spatial homogeneity. This work enhances the theoretical understanding of Bayesian additive regression trees under substantially relaxed smoothness assumptions. We provide a comprehensive study of asymptotic optimality and posterior contraction of Bayesian forests when the regression function has anisotropic smoothness that possibly varies over the function domain. The regression function can also be possibly discontinuous. We introduce a new class of sparse {\em piecewise heterogeneous anisotropic} Hölder functions and derive their minimax lower bound of estimation in high-dimensional scenarios under the $L_2$-loss. We then find that the Bayesian tree priors, coupled with a Dirichlet subset selection prior for sparse estimation in high-dimensional scenarios, adapt to unknown heterogeneous smoothness, discontinuity, and sparsity. These results show that Bayesian forests are uniquely suited for more general estimation problems that would render other default machine learning tools, such as Gaussian processes, suboptimal. Our numerical study shows that Bayesian forests often outperform other competitors such as random forests and deep neural networks, which are believed to work well for discontinuous or complicated smooth functions. Beyond nonparametric regression, we also examined posterior contraction of Bayesian forests for density estimation and binary classification using the technique developed in this study.

研究动机与目标

  • 通过在放宽的光滑性假设下分析其性能,填补贝叶斯森林理论上的空白,特别是针对各向异性和非齐次光滑性。
  • 解决现有极小极大渐近理论(通常基于各向同性或齐次性的限制性假设)与现实世界数据之间的脱节,后者通常表现出空间变化的光滑性。
  • 提出一种新函数类,以捕捉分段、异质、各向异性的光滑性,可能伴有不连续性,从而实现更真实的非参数函数估计。
  • 证明BART结合狄利克雷子集选择先验可在该一般函数类上实现后验收缩速率的极小极大最优性。
  • 将理论结果从回归推广至密度估计与二值分类,展示该框架的广泛适用性。

提出的方法

  • 引入一类新函数——分段异质各向异性霍尔德函数,其中每个矩形子区域具有独立的各向异性光滑度,光滑度可在不同区域间变化,并允许存在不连续性。
  • 为该新函数类在高维设定下推导$ L_2 $-损失下的极小极大下界,确立最优性的理论基准。
  • 通过在树划分上使用狄利克雷先验,并在回归系数上使用稀疏子集选择先验,以诱导稀疏性并适应未知的光滑度与不连续性。
  • 应用后验收缩理论,证明BART后验即使在光滑度空间变化且先验未知的情况下,仍能以极小极大速率集中。
  • 利用狄利克雷过程的棒折表示法,界定稀疏恢复的概率,确保先验偏好低维、相关特征。
  • 利用集中不等式与伽马函数的性质,特别是高维下小集中参数的情形,建立后验集中性的理论界。

实验结果

研究问题

  • RQ1在高维回归中,贝叶斯森林能否在非齐次、各向异性光滑性下实现极小极大最优估计速率?
  • RQ2当回归函数在输入空间不同区域表现出不连续性或变化的光滑性时,BART的表现如何?
  • RQ3能否证明BART的后验收缩速率与一类广义分段异质各向异性函数的极小极大下界一致?
  • RQ4树基先验与狄利克雷子集选择的结合,是否能实现对稀疏性、不连续性与异质光滑性的自适应估计,而无需先验知识?
  • RQ5理论结果在多大程度上可从回归推广至密度估计与二值分类?

主要发现

  • 论文为高维非参数回归中一类新的分段异质各向异性霍尔德函数(可能含有不连续性)建立了$ L_2 $-风险的极小极大下界。
  • 结合狄利克雷子集选择先验的BART实现了极小极大最优的后验收缩速率,即使在不同区域和方向的光滑度发生变化时亦然。
  • 后验概率落在真实函数$ \eta^* $的$ \epsilon $-邻域内的下界为$ \exp\{-C\xi s\log(p/\epsilon)\} $,证实了有效的集中性。
  • 先验通过将超过$ s $个变量活跃的配置的概率分配为指数级小值,确保了稀疏性,满足$ \Pi(\min_{|S|=s} \sum_{j\notin S} \eta_j \geq \epsilon) \leq \exp\{-C(\xi-1)s\log p - \log \epsilon\} $。
  • 理论结果表明,BART可自适应未知光滑度、不连续性与稀疏性,无需调参,在复杂场景中优于高斯过程及其他默认机器学习工具。
  • 数值实验表明,BART在估计不连续或高度异质光滑函数时,优于随机森林与深度神经网络,实证验证了理论结论。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。