Skip to main content
QUICK REVIEW

[論文レビュー] Optimal Learning via the Fourier Transform for Sums of Independent Integer Random Variables

Ilias Diakonikolas, Daniel M. Kane|arXiv (Cornell University)|May 4, 2015
Machine Learning and Algorithms参考文献 52被引用数 12
ひとこと要約

本稿では、フーリエ変換を用いて独立整数確率変数の和(SIIRV)を学習する計算効率の良いアルゴリズムを提示する。この手法は、$ widetilde{O}(k/\epsilon^2)$ の最適な標本複雑度を達成し、情報理論的下界 $ Theta((k/\epsilon^2)\sqrt{\log(1/\epsilon)})$ をも証明する。主な洞察は、$k$-SIIRV のフーリエ変換が近似的にスパースであることであり、これによりスペクトル的手法を用いた正確で効率的な学習が可能になる。

ABSTRACT

We study the structure and learnability of sums of independent integer random variables (SIIRVs). For $k \in \mathbb{Z}_{+}$, a $k$-SIIRV of order $n \in \mathbb{Z}_{+}$ is the probability distribution of the sum of $n$ independent random variables each supported on $\{0, 1, \dots, k-1\}$. We denote by ${\cal S}_{n,k}$ the set of all $k$-SIIRVs of order $n$. In this paper, we tightly characterize the sample and computational complexity of learning $k$-SIIRVs. More precisely, we design a computationally efficient algorithm that uses $\widetilde{O}(k/ε^2)$ samples, and learns an arbitrary $k$-SIIRV within error $ε,$ in total variation distance. Moreover, we show that the {\em optimal} sample complexity of this learning problem is $Θ((k/ε^2)\sqrt{\log(1/ε)}).$ Our algorithm proceeds by learning the Fourier transform of the target $k$-SIIRV in its effective support. Its correctness relies on the {\em approximate sparsity} of the Fourier transform of $k$-SIIRVs -- a structural property that we establish, roughly stating that the Fourier transform of $k$-SIIRVs has small magnitude outside a small set. Along the way we prove several new structural results about $k$-SIIRVs. As one of our main structural contributions, we give an efficient algorithm to construct a sparse {\em proper} $ε$-cover for ${\cal S}_{n,k},$ in total variation distance. We also obtain a novel geometric characterization of the space of $k$-SIIRVs. Our characterization allows us to prove a tight lower bound on the size of $ε$-covers for ${\cal S}_{n,k}$, and is the key ingredient in our tight sample complexity lower bound. Our approach of exploiting the sparsity of the Fourier transform in distribution learning is general, and has recently found additional applications.

研究の動機と目的

  • 任意の $k$-SIIRV を $\epsilon$ 総変動距離内で学習するための標本複雑度と計算複雑度を同定すること。
  • $k$-SIIRV の学習における標本複雑度のタイトな情報理論的上界と下界を確立すること。
  • 標本数 $ widetilde{O}(k/\epsilon^2)$ を用いる計算的に効率的な学習アルゴリズムを開発すること。
  • 新規な構造的性質の証明:$k$-SIIRV のフーリエ変換は、その有効なサポートにおいて近似的にスパースである。
  • 総変動距離における $k$-SIIRV の空間に対する効率的なスパース $\epsilon$-カバーを構築すること。

提案手法

  • アルゴリズムは、ターゲット $k$-SIIRV のフーリエ変換をその有効サポート内で学習し、変換の近似的なスパarsityを活用する。
  • $k$-SIIRV のフーリエ係数が小さな集合の外では急速に減少することを活用し、少ない標本数で効率的に推定可能である。
  • 分布空間の幾何学的およびスペクトル的性質を用いて、$\mathcal{S}_{n,k}$ に対するスパースな適切な $\epsilon$-カバーを構築する。
  • $k$-SIIRV の新しい幾何的特徴付けが導出され、$\epsilon$-カバーのサイズに対するタイトな境界を可能にする。
  • 下界の証明には、適切に構築された $k$-SIIRV のハイパーキューブ上でのアッソアドの補題が用いられ、制御された総変動距離が保証される。
  • 解析は、ポアソン二項分布およびその条件付き分布の集中と反集中の不等式に依存する。

実験結果

リサーチクエスチョン

  • RQ1総変動距離 $\epsilon$ 以内で $k$-SIIRV を学習するための最適な標本複雑度は何か?
  • RQ2$k$-SIIRV のフーリエ変換を用いて、計算的に効率的な学習アルゴリズムを設計できるか?
  • RQ3$k$-SIIRV のフーリエ変換は近似的にスパースか? そして、この性質はどのように学習に活用できるか?
  • RQ4総変動距離における $k$-SIIRV の空間に対する $\epsilon$-カバーの最小サイズは何か?
  • RQ5情報理論的手法を用いて、標本複雑度のタイトな下界を確立できるか?

主な発見

  • 総変動距離 $\epsilon$ 以内で $k$-SIIRV を学習するための最適な標本複雑度は $\Theta\left(\frac{k}{\epsilon^2}\sqrt{\log(1/\epsilon)}\right)$ である。
  • 提案されたアルゴリズムは、この境界を $ widetilde{O}(k/\epsilon^2)$ の標本数で達成し、標本サイズに多項式時間で実行される。
  • 任意の $k$-SIIRV のフーリエ変換は、$L^2$ マスの大部分が $O(k/\epsilon^2)$ のサイズの集合に集中するという意味で、近似的にスパースである。
  • サイズ $ widetilde{O}(k/\epsilon^2)$ の $\mathcal{S}_{n,k}$ に対するスパースな適切な $\epsilon$-カバーを効率的に構築可能である。
  • $k$-SIIRV の幾何的特徴付けにより、任意の $\epsilon$-カバーのサイズに対するタイトな下界を証明でき、カバー構築の最適性を確認できる。
  • 下界構築では、対ごとの総変動距離が $\Omega(2^{-Cn})$ であるような $k$-SIIRV のハイパーキューブが用いられ、標本複雑度の下界が $\Omega\left((k/\epsilon^2)\sqrt{\log(1/\epsilon)}\right)$ となる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。