Skip to main content
QUICK REVIEW

[论文解读] Optimal Learning via the Fourier Transform for Sums of Independent Integer Random Variables

Ilias Diakonikolas, Daniel M. Kane|arXiv (Cornell University)|May 4, 2015
Machine Learning and Algorithms参考文献 52被引用 12
一句话总结

本文提出了一种计算高效的算法,通过傅里叶变换学习独立整数值随机变量之和(SIIRVs),实现了最优的样本复杂度 $\widetilde{O}(k/\epsilon^2)$,并证明了信息论下界 $\Theta((k/\epsilon^2)\sqrt{\log(1/\epsilon)})$ 的匹配性。关键洞见在于 $k$-SIIRVs 的傅里叶变换具有近似稀疏性,从而可通过谱方法实现精确且高效的学习。

ABSTRACT

We study the structure and learnability of sums of independent integer random variables (SIIRVs). For $k \in \mathbb{Z}_{+}$, a $k$-SIIRV of order $n \in \mathbb{Z}_{+}$ is the probability distribution of the sum of $n$ independent random variables each supported on $\{0, 1, \dots, k-1\}$. We denote by ${\cal S}_{n,k}$ the set of all $k$-SIIRVs of order $n$. In this paper, we tightly characterize the sample and computational complexity of learning $k$-SIIRVs. More precisely, we design a computationally efficient algorithm that uses $\widetilde{O}(k/ε^2)$ samples, and learns an arbitrary $k$-SIIRV within error $ε,$ in total variation distance. Moreover, we show that the {\em optimal} sample complexity of this learning problem is $Θ((k/ε^2)\sqrt{\log(1/ε)}).$ Our algorithm proceeds by learning the Fourier transform of the target $k$-SIIRV in its effective support. Its correctness relies on the {\em approximate sparsity} of the Fourier transform of $k$-SIIRVs -- a structural property that we establish, roughly stating that the Fourier transform of $k$-SIIRVs has small magnitude outside a small set. Along the way we prove several new structural results about $k$-SIIRVs. As one of our main structural contributions, we give an efficient algorithm to construct a sparse {\em proper} $ε$-cover for ${\cal S}_{n,k},$ in total variation distance. We also obtain a novel geometric characterization of the space of $k$-SIIRVs. Our characterization allows us to prove a tight lower bound on the size of $ε$-covers for ${\cal S}_{n,k}$, and is the key ingredient in our tight sample complexity lower bound. Our approach of exploiting the sparsity of the Fourier transform in distribution learning is general, and has recently found additional applications.

研究动机与目标

  • 刻画在总变差距离 $\epsilon$ 内学习任意 $k$-SIIRV 的样本复杂度与计算复杂度。
  • 建立学习 $k$-SIIRVs 的样本复杂度的紧致信息论上下界。
  • 设计一种计算高效的算法,使用 $\widetilde{O}(k/\epsilon^2)$ 个样本。
  • 证明一个新颖的结构性质:$k$-SIIRVs 的傅里叶变换在其有效支撑上近似稀疏。
  • 在总变差距离下,为 $k$-SIIRVs 的空间构建一个高效的稀疏 $\epsilon$-覆盖。

提出的方法

  • 该算法在有效支撑内学习目标 $k$-SIIRV 的傅里叶变换,利用了变换的近似稀疏性。
  • 它利用了 $k$-SIIRVs 的傅里叶系数在小集合之外迅速衰减的特性,从而可用少量样本实现高效估计。
  • 该方法利用分布空间的几何与谱性质,构建了 $\mathcal{S}_{n,k}$ 的稀疏恰当 $\epsilon$-覆盖。
  • 推导出 $k$-SIIRVs 的新颖几何表征,从而实现对 $\epsilon$-覆盖大小的紧致界。
  • 下界证明使用了在精心构造的 $k$-SIIRVs 超立方体上应用 Assouad 引理,其总变差距离受控。
  • 分析依赖于泊松二项分布及其条件分布的集中与反集中不等式。

实验结果

研究问题

  • RQ1在总变差距离 $\epsilon$ 内学习 $k$-SIIRV 的最优样本复杂度是什么?
  • RQ2能否利用 $k$-SIIRVs 的傅里叶变换设计计算高效的算法?
  • RQ3$k$-SIIRVs 的傅里叶变换是否近似稀疏,以及如何利用这一性质进行学习?
  • RQ4在总变差距离下,$k$-SIIRVs 空间的最小 $\epsilon$-覆盖大小是多少?
  • RQ5能否使用信息论方法建立样本复杂度的紧致下界?

主要发现

  • 在总变差距离 $\epsilon$ 内学习 $k$-SIIRV 的最优样本复杂度为 $\Theta\left(\frac{k}{\epsilon^2}\sqrt{\log(1/\epsilon)}\right)$。
  • 所提出的算法以 $\widetilde{O}(k/\epsilon^2)$ 个样本达到该界,且运行时间在样本大小的多项式时间内。
  • 任意 $k$-SIIRV 的傅里叶变换近似稀疏,其大部分 $L^2$ 质量集中于大小为 $O(k/\epsilon^2)$ 的集合中。
  • 可高效构造大小为 $\widetilde{O}(k/\epsilon^2)$ 的 $\mathcal{S}_{n,k}$ 的稀疏恰当 $\epsilon$-覆盖。
  • $k$-SIIRVs 的几何表征使得能够证明任何 $\epsilon$-覆盖大小的紧致下界,从而确认了覆盖构造的最优性。
  • 下界构造使用了一个 $k$-SIIRVs 的超立方体,其两两之间的总变差距离为 $\Omega(2^{-Cn})$,从而导出样本复杂度下界 $\Omega\left((k/\epsilon^2)\sqrt{\log(1/\epsilon)}\right)$。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。