Skip to main content
QUICK REVIEW

[论文解读] Optimality of Maximum Likelihood for Log-Concave Density Estimation and Bounded Convex Regression

Gil Kur, Yuval Dagan|arXiv (Cornell University)|Mar 13, 2019
Statistical Methods and Inference参考文献 54被引用 18
一句话总结

本文建立了高维情况下对数凹密度估计和有界凸回归的最大似然估计(MLE)的极小极大最优性,表明MLE在平方Hellinger距离和$L_2$距离下达到$\Theta_d(n^{-2/(d+1)})$的收敛速率,其中$d \geq 4$,将先前仅适用于$d \leq 3$的最优性结果扩展至更高维。此外,本文在总变差距离下推导出一个精确的$\Theta_d(n^{-2/(d+4)})$极小极大速率,并证明估计对数凹密度需要指数级增长的样本量,从而改进了已知下界中的维数常数。

ABSTRACT

In this paper, we study two problems: (1) estimation of a $d$-dimensional log-concave distribution and (2) bounded multivariate convex regression with random design with an underlying log-concave density or a compactly supported distribution with a continuous density. First, we show that for all $d \ge 4$ the maximum likelihood estimators of both problems achieve an optimal risk of $Θ_d(n^{-2/(d+1)})$ (up to a logarithmic factor) in terms of squared Hellinger distance and $L_2$ squared distance, respectively. Previously, the optimality of both these estimators was known only for $d\le 3$. We also prove that the $ε$-entropy numbers of the two aforementioned families are equal up to logarithmic factors. We complement these results by proving a sharp bound $Θ_d(n^{-2/(d+4)})$ on the minimax rate (up to logarithmic factors) with respect to the total variation distance. Finally, we prove that estimating a log-concave density - even a uniform distribution on a convex set - up to a fixed accuracy requires the number of samples \emph{at least} exponential in the dimension. We do that by improving the dimensional constant in the best known lower bound for the minimax rate from $2^{-d}\cdot n^{-2/(d+1)}$ to $c\cdot n^{-2/(d+1)}$ (when $d\geq 2$).

研究动机与目标

  • 建立高维情形下($d \geq 4$)对数凹密度估计和有界凸回归的最大似然估计(MLE)的极小极大最优性,将已知结果从$d \leq 3$扩展至更高维。
  • 推导在总变差距离下的精确极小极大速率,表明速率在对数因子内为$\Theta_d(n^{-2/(d+4)})$。
  • 证明估计对数凹密度——即使是在凸集上的均匀分布——也需要指数级增长的样本复杂度,改进了已知下界中的维数常数。

提出的方法

  • 使用仿射变换对数据进行标准化,利用前$n/3$个样本的经验协方差矩阵,将任意对数凹分布转化为近似各向同性的分布。
  • 对标准化后的样本应用MLE估计器,利用已知的各向同性对数凹密度的最优性结果。
  • 采用三阶段采样策略:$n/3$个样本用于协方差估计,$n/3$个样本用于均值估计,$n/3$个样本用于变换后数据的密度估计。
  • 利用平方凸函数的尾部公式和集中不等式,控制对数凹测度上的经验过程偏差。
  • 应用 bracketing entropy 和熵数技术,将函数类的复杂度与估计风险联系起来。
  • 基于熵积分的不动点论证,建立非Donsker情形下MLE的最优性,将先前结果从Donsker条件扩展至更一般情形。

实验结果

研究问题

  • RQ1在维度$d \geq 4$时,最大似然估计器对对数凹密度估计是否最优?
  • RQ2对具有对数凹或紧支集密度的有界多元凸回归,估计的极小极大速率是什么?
  • RQ3MLE在总变差距离下是否达到最优速率,其对维度和样本量的依赖关系如何?
  • RQ4估计对数凹密度的样本复杂度能否被维数的指数函数从下方界定,且能否改进已知下界中的常数?
  • RQ5对数凹密度类和有界凸回归类的熵数在对数因子内是否等价?

主要发现

  • 对于$d$维对数凹密度估计,MLE在平方Hellinger距离下的极小极大风险为$\Theta_d(n^{-2/(d+1)})$,对所有$d \geq 4$成立,将最优性结果从$d \leq 3$扩展至更高维。
  • 对于具有对数凹或紧支集密度的有界多元凸回归,MLE在$L_2$平方距离下也达到相同的$\Theta_d(n^{-2/(d+1)})$速率。
  • 对数凹密度类和有界凸回归类的$\epsilon$-熵数在对数因子内相等。
  • 在总变差距离下,建立了精确的极小极大速率$\Theta_d(n^{-2/(d+4)})$,在对数因子内成立。
  • 本文证明,以固定精度估计对数凹密度至少需要$\Omega(c^n)$个样本,其中$c > 0$,将先前的下界从$2^{-d} \cdot n^{-2/(d+1)}$改进为$c \cdot n^{-2/(d+1)}$($d \geq 2$),表明其对维数具有指数依赖性。
  • 分析确认,在非Donsker情形下,MLE对这些类是最优的,通过基于熵积分的不动点论证,将结果扩展至经典Donsker型假设之外。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。