[论文解读] Fast rates for empirical risk minimization over càdlàg functions with bounded sectional variation norm
本文在高维设定下建立了对有界截面变差范数的 càdlàg 函数类上经验风险最小化的快速收敛速率。通过利用这类函数的新型表示方法并控制其 bracketing 熵,作者推导出收敛速率为 $ O_P(n^{-1/3}( ext{log}ackslash,n)^{2(d-1)/3}a_n) $,其中维度 $ d $ 仅通过对数因子影响速率,使该方法在次高斯误差下特别适用于高维非参数估计。
Empirical risk minimization over classes functions that are bounded for some version of the variation norm has a long history, starting with Total Variation Denoising (Rudin et al., 1992), and has been considered by several recent articles, in particular Fang et al., 2019 and van der Laan, 2015. In this article, we consider empirical risk minimization over the class $\mathcal{F}_d$ of càdlàg functions over $[0,1]^d$ with bounded sectional variation norm (also called Hardy-Krause variation). We show how a certain representation of functions in $\mathcal{F}_d$ allows to bound the bracketing entropy of sieves of $\mathcal{F}_d$, and therefore derive rates of convergence in nonparametric function estimation. Specifically, for sieves whose growth is controlled by some rate $a_n$, we show that the empirical risk minimizer has rate of convergence $O_P(n^{-1/3} (\log n)^{2(d-1)/3} a_n)$. Remarkably, the dimension only affects the rate in $n$ through the logarithmic factor, making this method especially appropriate for high dimensional problems. In particular, we show that in the case of nonparametric regression over sieves of càdlàg functions with bounded sectional variation norm, this upper bound on the rate of convergence holds for least-squares estimators, under the random design, sub-exponential errors setting.
研究动机与目标
- 在高维设定下,建立对有界截面变差范数的 càdlàg 函数类上经验风险最小化(ERM)的快速收敛速率。
- 刻画该函数类中筛子的 bracketing 熵,以实现非渐近风险界。
- 证明收敛速率仅通过一个对数因子受维度影响,使该方法在高维下具有可扩展性。
- 将 ERM 的适用性扩展至一般设计下具有次高斯误差假设的非参数回归与逻辑回归。
提出的方法
- 作者利用有界变差测度之和表示有界截面变差范数的 càdlàg 函数,实现对熵的精确控制。
- 他们推导出函数类 $ \mathcal{F}_d $ 中筛子的 bracketing 熵界,这是推导收敛速率的核心。
- 应用剥皮技术控制熵积分,从而推导出最终的收敛速率。
- 关键步骤是将参数空间的 $ L^2 $-norm bracketing 与损失函数的 Bernstein 范数联系起来,从而可应用集中不等式。
- 该方法应用于随机设计下具有次高斯误差的最小二乘估计,表明 ERM 估计器可达到所推导的速率。
- 理论结果通过将 ERM 估计器表示为基于基函数的 LASSO 类问题得到支持,确保了计算上的可行性。
实验结果
研究问题
- RQ1在高维设定下,能否为有界截面变差范数的 càdlàg 函数类上的经验风险最小化建立快速收敛速率?
- RQ2在这样的非参数估计问题中,维度 $ d $ 如何影响收敛速率?
- RQ3所推导的速率是否可在一般损失函数下实现,包括单峰利普希茨损失,以及在次高斯误差模型下?
- RQ4此类函数类上的经验风险最小化器是否计算上可行,能否重述为 LASSO 问题?
- RQ5在具有次高斯误差的非参数回归中,该速率 $ n^{-1/3}( ext{log}\,n)^{2(d-1)/3}a_n $ 是否对最小二乘估计器成立?
主要发现
- 对有界截面变差范数的 càdlàg 函数类的筛子上,经验风险最小化器实现了 $ O_P(n^{-1/3}( ext{log}\,n)^{2(d-1)/3}a_n) $ 的收敛速率,其中 $ a_n $ 控制筛子的增长。
- 维度 $ d $ 仅通过一个对数因子影响速率,使该方法在高维问题中极具适用性。
- 通过将 càdlàg 函数表示为有界变差测度之和,对函数类的 bracketing 熵进行了有界控制,从而支持基于熵的收敛性分析。
- 该速率在随机设计和次高斯误差下对非参数回归的最小二乘估计成立,且无需假设高斯或格点设计。
- ERM 估计器可表示为具有 $ O((ne/d)^d) $ 个基函数的 LASSO 问题,确保了计算上的可行性。
- 结果可扩展至单峰利普希茨损失,涵盖重要情形如响应有界的非参数最小二乘估计和逻辑回归。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。