[论文解读] Convex Regression in Multidimensions: Suboptimality of Least Squares Estimators
该论文证明,在维度 $ d \geq 5 $ 时,多维凸回归中的最小二乘估计量(LSE)是次优的,其风险量级为 $ n^{-2/d} $,而极小极大风险为 $ n^{-4/(d+4)} $。作者首次建立了在多面体域上完整凸LSE的最坏情况和自适应收敛速率,并推导出凸函数类的新度量熵界,表明尽管LSE无需调参且是一致的,但在高维情形下仍无法达到最优极小极大速率。
Under the usual nonparametric regression model with Gaussian errors, Least Squares Estimators (LSEs) over natural subclasses of convex functions are shown to be suboptimal for estimating a $d$-dimensional convex function in squared error loss when the dimension $d$ is 5 or larger. The specific function classes considered include: (i) bounded convex functions supported on a polytope (in random design), (ii) Lipschitz convex functions supported on any convex domain (in random design), (iii) convex functions supported on a polytope (in fixed design). For each of these classes, the risk of the LSE is proved to be of the order $n^{-2/d}$ (up to logarithmic factors) while the minimax risk is $n^{-4/(d+4)}$, when $d \ge 5$. In addition, the first rate of convergence results (worst case and adaptive) for the unrestricted convex LSE are established in fixed-design for polytopal domains for all $d \geq 1$. Some new metric entropy results for convex functions are also proved which are of independent interest.
研究动机与目标
- 研究在具有高斯误差的多维非参数回归中,凸最小二乘估计量(LSE)的理论最优性。
- 确定LSE在维度 $ d \geq 5 $ 时是否能达到凸回归的极小极大收敛速率。
- 为所有 $ d \geq 1 $ 的多面体域上的完整凸LSE,首次建立最坏情况和自适应收敛速率。
- 推导凸函数类的新度量熵界,这些结果在经验过程理论中具有独立兴趣。
提出的方法
- 作者在三种设定下分析了凸LSE的风险:(i) 在固定设计下,多面体上的凸函数;(ii) 在随机设计下,多面体上的有界凸函数;(iii) 在随机设计下,任意凸域上的Lipschitz凸函数。
- 他们采用度量熵技术,包括Dudley的熵积分和Sudakov的下界估计,推导出凸函数类熵的下界。
- 关键构造涉及基于 $ \Omega \subseteq \mathbb{R}^d $ 中点网格的一族分段仿射凸函数,其与真实函数 $ f_0 $ 的 $ L^2 $-距离受控。
- 次优性证明依赖于一个基于有限函数集 $ \{f_1, \dots, f_N\} $ 的检验方法,该集合中函数对之间的 $ \ell_{{\mathbb{P}}_n} $-距离有下界,且一致逼近误差有上界。
- 应用Hoeffding不等式控制经验 $ \ell_2 $-损失与真实 $ \ell_2 $-损失之间的偏差,从而构建一个检验框架,用于下界估计LSE的风险。
- 作者推导出在 $ d $ 维多面体上 $ L $-Lipschitz凸函数类的新度量熵界,表明熵复杂度为 $ \epsilon^{-d/2} $ 量级。
实验结果
研究问题
- RQ1在维度 $ d \geq 5 $ 时,凸最小二乘估计量(LSE)在多维凸回归中是否为极小极大最优?
- RQ2在所有 $ d \geq 1 $ 的多面体域上,完整凸LSE的收敛速率是多少?
- RQ3在高维凸回归中,LSE的风险能否以极小极大风险为下界?
- RQ4在 $ d $ 维凸域上,$ L $-Lipschitz凸函数类的度量熵是多少?
- RQ5尽管LSE无需调参且是一致的,它是否仍无法达到极小极大速率?
主要发现
- 当 $ d \geq 5 $ 时,凸LSE的风险为 $ \Omega(n^{-2/d}) $,而极小极大风险为 $ O(n^{-4/(d+4)}) $,证明了LSE在平方误差损失下的次优性。
- 在所有 $ d \geq 1 $ 的多面体域上,凸LSE的最坏情况收敛速率为 $ n^{-2/d} $,对数因子除外。
- 凸LSE的自适应收敛速率也被确定为 $ n^{-2/d} $,与最坏情况速率一致。
- 作者证明了在 $ d $ 维多面体上 $ L $-Lipschitz凸函数类的新度量熵界,表明熵复杂度为 $ \epsilon^{-d/2} $,该结果在对数因子意义下是紧的。
- LSE在三种设定下均被证明是次优的:(i) 多面体上的凸函数(固定设计);(ii) 多面体上的有界凸函数(随机设计);(iii) 任意凸域上的Lipschitz凸函数(随机设计)。
- 证明技术依赖于构造一个大小为 $ N \sim \epsilon^{-d/2} $ 的有限函数集,其两两之间的 $ \ell_2 $-距离 $ \geq \epsilon $,且一致逼近误差 $ \leq 4\epsilon $,从而通过检验方法实现下界估计。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。