[論文レビュー] Learning Coverage Functions and Private Release of Marginals
本稿では、PMACモデルにおけるカバレッジ関数の学習のための最初の完全多項式時間アルゴリズムを提示する。このアルゴリズムは、入力の $1-\delta$ 分の1で、乗法的近似誤差が $1+\gamma$ 以内に収まるように、時間 $\mathsf{poly}(n,1/\gamma,1/\delta)$ で実行される。さらに、この学習フレームワークを用いて、低平均誤差で $k$-ウェイマージナルを微分プライベートにリリースすることが可能になる。
We study the problem of approximating and learning coverage functions. A function $c: 2^{[n]} ightarrow \mathbf{R}^{+}$ is a coverage function, if there exists a universe $U$ with non-negative weights $w(u)$ for each $u \in U$ and subsets $A_1, A_2, \ldots, A_n$ of $U$ such that $c(S) = \sum_{u \in \cup_{i \in S} A_i} w(u)$. Alternatively, coverage functions can be described as non-negative linear combinations of monotone disjunctions. They are a natural subclass of submodular functions and arise in a number of applications. We give an algorithm that for any $γ,δ>0$, given random and uniform examples of an unknown coverage function $c$, finds a function $h$ that approximates $c$ within factor $1+γ$ on all but $δ$-fraction of the points in time $poly(n,1/γ,1/δ)$. This is the first fully-polynomial algorithm for learning an interesting class of functions in the demanding PMAC model of Balcan and Harvey (2011). Our algorithms are based on several new structural properties of coverage functions. Using the results in (Feldman and Kothari, 2014), we also show that coverage functions are learnable agnostically with excess $\ell_1$-error $ε$ over all product and symmetric distributions in time $n^{\log(1/ε)}$. In contrast, we show that, without assumptions on the distribution, learning coverage functions is at least as hard as learning polynomial-size disjoint DNF formulas, a class of functions for which the best known algorithm runs in time $2^{ ilde{O}(n^{1/3})}$ (Klivans and Servedio, 2004). As an application of our learning results, we give simple differentially-private algorithms for releasing monotone conjunction counting queries with low average error. In particular, for any $k \leq n$, we obtain private release of $k$-way marginals with average error $\barα$ in time $n^{O(\log(1/\barα))}$.
研究の動機と目的
- PMACモデルにおいて、仮説が大多数の入力に対して乗法的近似を行うように、カバレッジ関数のための効率的学習アルゴリズムを開発すること。
- 要求の厳しいPMACモデルにおいて、非自明なサブモジュラ関数クラスの最初の多項式時間学習アルゴリズムを確立すること。
- 学習されたカバレッジ関数を活用して、低平均誤差で単調な論理積カウントクエリの微分プライベートなリリースを可能にすること。
- 一様分布、積分布、対称分布を含むさまざまな分布におけるカバレッジ関数の学習可能性を調査すること。
- 分布に依存しないカバレッジ関数の学習の計算複雑性を、多項式サイズの互いに素なDNF式の学習という難しい問題と関連づけ、明確化すること。
提案手法
- カバレッジ関数を単調論理和の非負の線形結合として表す、新しい構造的特徴づけを活用する。
- 各閾値関数を単調論理和の学習に還元することで、閾値に基づく近似スキームを採用する。
- 多項式回帰と変数同定技術を用いて、カバレッジ関数分解における単調論理和の係数を推定する。
- ブースティング風の還元を適用:ターゲット関数の閾値関数を学習し、それらを組み合わせて元の関数を再構築し、$\ell_1$ 誤差を制御する。
- $\ell_1$ 誤差によるカバレッジ関数の学習を、非負の単調論理和の和の閾値関数のPAC学習に還元する。
- パrameterized閾値処理機構を用いて近似誤差を制限し、最終的な仮説がターゲット関数に対して $\ell_1$ 誤差が $\epsilon$ 以内に収まるように保証する。
実験結果
リサーチクエスチョン
- RQ1分布に関する仮定なしに、PMACモデルにおけるカバレッジ関数を効率的に学習することは可能か?
- RQ2一般の分布におけるカバレッジ関数の学習の計算複雑性は何か? また、これと既知の難易度の高いクラス(例えば、互いに素なDNF式)との関係は?
- RQ3この学習フレームワークを拡張して、低平均誤差で $k$-ウェイマージナルの微分プライベートリリースを可能にすることができるか?
- RQ4積分布および対称分布上でのカバレッジ関数の効率的 $\ell_1$-誤差学習は可能か?
- RQ5カバレッジ関数のどの構造的性質が、PMACおよび $\ell_1$-誤差モデルにおけるその効率的学習を可能にするのか?
主な発見
- 本稿では、PMACモデルにおけるカバレッジ関数の学習のための最初の完全多項式時間アルゴリズムを提示する。このアルゴリズムは、時間 $\mathsf{poly}(n,1/\gamma,1/\delta)$ で、入力の $1-\delta$ 分の1において $1+\gamma$ の乗法的近似誤差を達成する。
- 積分布および対称分布上でのカバレッジ関数の $\ell_1$-誤差学習が、超過誤差 $\epsilon$ に対して時間 $n^{\log(1/\epsilon)}$ で達成可能であることが示された。
- 分布に依存しないカバレッジ関数の学習は、多項式サイズの互いに素なDNF式のPAC学習ほど難しい。このクラスに対する最良の既知のアルゴリズムは時間 $2^{\tilde{O}(n^{1/3})}$ で実行される。
- このフレームワークにより、平均誤差 $\bar{\alpha}$ で $k$-ウェイマージナルを微分プライベートにリリースすることができ、実行時間は $n^{O(\log(1/\bar{\alpha}))}$ である。
- $\ell_1$-誤差によるカバレッジ関数の学習が、任意の固定分布上では、非負の単調論理和の和の閾値関数のPAC学習に還元可能であることを示す還元が確立された。
- 閾値近似の組み合わせと、閾値処理プロセスによる誤差伝搬の制限により、アルゴリズムは $\ell_1$ 誤差を $\mathop{\mathbf{E}}[|y - c^*(x)|] + 3\epsilon$ 以内に抑えることができる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。