Skip to main content
QUICK REVIEW

[论文解读] Sample complexity of population recovery

Yury Polyanskiy, Ananda Theertha Suresh|arXiv (Cornell University)|Feb 18, 2017
Bayesian Methods and Mixture Models参考文献 26被引用 7
一句话总结

本文在两种模型下建立了总体恢复的最优样本复杂度:丢失(擦除)模型和噪声(对称误差)模型。结果表明,在丢失恢复中,样本复杂度为 $\tilde{\Theta}(\delta^{-2\max\{\epsilon/(1-\epsilon),1\}})$,在 $\epsilon = 1/2$ 处出现相变;而在噪声恢复中,其复杂度为 $\exp(\Theta(d^{1/3}\log^{2/3}(1/\delta)))$,该结果通过求解线性规划后进行经验均值估计得到,其对偶问题对应于 Le Cam 的极小极大下界。

ABSTRACT

The problem of population recovery refers to estimating a distribution based on incomplete or corrupted samples. Consider a random poll of sample size $n$ conducted on a population of individuals, where each pollee is asked to answer $d$ binary questions. We consider one of the two polling impediments: (a) in lossy population recovery, a pollee may skip each question with probability $ε$, (b) in noisy population recovery, a pollee may lie on each question with probability $ε$. Given $n$ lossy or noisy samples, the goal is to estimate the probabilities of all $2^d$ binary vectors simultaneously within accuracy $δ$ with high probability. This paper settles the sample complexity of population recovery. For lossy model, the optimal sample complexity is $ ildeΘ(δ^{-2\max\{\fracε{1-ε},1\}})$, improving the state of the art by Moitra and Saks in several ways: a lower bound is established, the upper bound is improved and the result depends at most on the logarithm of the dimension. Surprisingly, the sample complexity undergoes a phase transition from parametric to nonparametric rate when $ε$ exceeds $1/2$. For noisy population recovery, the sharp sample complexity turns out to be more sensitive to dimension and scales as $\exp(Θ(d^{1/3} \log^{2/3}(1/δ)))$ except for the trivial cases of $ε=0,1/2$ or $1$. For both models, our estimators simply compute the empirical mean of a certain function, which is found by pre-solving a linear program (LP). Curiously, the dual LP can be understood as Le Cam's method for lower-bounding the minimax risk, thus establishing the statistical optimality of the proposed estimators. The value of the LP is determined by complex-analytic methods.

研究动机与目标

  • 确定从被污染或不完整的样本中估计高维分布的基本统计极限。
  • 解决丢失和噪声总体恢复模型下的样本复杂度问题,这两类模型在从部分或错误观测中学习中具有核心地位。
  • 通过证明个体恢复与总体恢复在样本复杂度上至多对数因子等价,统一二者分析。
  • 利用复分析和极小极大理论(特别是 Le Cam 方法)的工具,推导出精确的上下界。
  • 证明最优估计器是通过求解线性规划得到的精心构造函数的经验均值。

提出的方法

  • 所提出的估计器计算通过求解线性规划(LP)得到的函数的经验均值,该 LP 事先计算以确保统计最优性。
  • LP 的对偶被解释为 Le Cam 方法用于推导极小极大下界,从而确立估计器的统计最优性。
  • 使用复分析技术,包括 $H^\infty$-范数边界和生成函数的泰勒级数分析,以评估 LP 的值。
  • 对于丢失模型,分析利用 Parseval 恒等式和柯西-施瓦茨不等式,以界定与恢复误差相关的生成函数的 $A$-范数。
  • 对于噪声模型,方法利用包含拉盖尔多项式的生成函数及渐近边界,以控制 $A$-范数并推导出对 $d$ 的指数依赖关系。
  • 分析表明,样本复杂度在 $\epsilon = 1/2$ 处发生相变,从参数速率转变为非参数速率。

实验结果

研究问题

  • RQ1在丢失(擦除)模型下,总体恢复的最优样本复杂度是多少?其如何依赖于擦除概率 $\epsilon$ 和精度 $\delta$?
  • RQ2噪声总体恢复的样本复杂度如何随维度 $d$ 和精度 $\delta$ 变化?其精确依赖关系是什么?
  • RQ3在适当的函数变换后,经验均值估计器是否能在两种模型中均达到极小极大最优速率?
  • RQ4对偶 LP 在建立极小极大下界中起什么作用?其与 Le Cam 方法有何关联?
  • RQ5为何丢失模型在 $\epsilon = 1/2$ 处出现相变?这如何影响收敛速率?

主要发现

  • 在丢失总体恢复模型中,最优样本复杂度为 $\tilde{\Theta}(\delta^{-2\max{\{\epsilon/(1-\epsilon),1\}}})$,相比先前工作,该结果通过建立紧致下界,将对维度的依赖降低至对数阶。
  • 在 $\epsilon = 1/2$ 处发生相变:当 $\epsilon < 1/2$ 时,速率为参数型($\delta^{-2}$);当 $\epsilon > 1/2$ 时,速率为非参数型($\delta^{-2\epsilon/(1-\epsilon)}$)。
  • 在噪声总体恢复模型中,最优样本复杂度为 $\exp(\Theta(d^{1/3}\log^{2/3}(1/\delta)))$,其对维度的敏感性远高于丢失情形。
  • 所提出的估计器是通过求解线性规划得到的函数的经验均值,其最优性由与 Le Cam 极小极大下界的对偶性所证明。
  • LP 的值通过复分析方法确定,包括生成函数 $H^\infty$-范数的边界估计及拉盖尔多项式的渐近分析。
  • 分析表明,两种模型的样本复杂度在丢失情况下(对数因子内)与维度 $d$ 无关,而在噪声情况下则随 $d^{1/3}$ 呈指数增长。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。