Skip to main content
QUICK REVIEW

[論文レビュー] Comparing distributions: $\ell_1$ geometry improves kernel two-sample testing

Meyer Scetbon, Gaël Varoquaux|arXiv (Cornell University)|Sep 19, 2019
Gaussian Processes and Bayesian Inference被引用数 5
ひとこと要約

本稿は、空間的または周波数領域の位置における解析関数の期待値の差の $\ell_1$ ノルムを用いる $\iota_1$-幾何に基づく二標本検定を提案する。この手法は、カーネル二標本検定を改善し、より高い検出力と高速な計算を達成する。また、解釈可能な特徴選択と、合成データおよび実世界のデータにおいて一貫した性能を示す。

ABSTRACT

Are two sets of observations drawn from the same distribution? This problem is a two-sample test. Kernel methods lead to many appealing properties. Indeed state-of-the-art approaches use the $L^2$ distance between kernel-based distribution representatives to derive their test statistics. Here, we show that $L^p$ distances (with $p\geq 1$) between these distribution representatives give metrics on the space of distributions that are well-behaved to detect differences between distributions as they metrize the weak convergence. Moreover, for analytic kernels, we show that the $L^1$ geometry gives improved testing power for scalable computational procedures. Specifically, we derive a finite dimensional approximation of the metric given as the $\ell_1$ norm of a vector which captures differences of expectations of analytic functions evaluated at spatial locations or frequencies (i.e, features). The features can be chosen to maximize the differences of the distributions and give interpretable indications of how they differs. Using an $\ell_1$ norm gives better detection because differences between representatives are dense as we use analytic kernels (non-zero almost everywhere). The tests are consistent, while much faster than state-of-the-art quadratic-time kernel-based tests. Experiments on artificial and real-world problems demonstrate improved power/time tradeoff than the state of the art, based on $\ell_2$ norms, and in some cases, better outright power than even the most expensive quadratic-time tests.

研究の動機と目的

  • 既存のカーネル二標本検定が $\ell_2$-ベースの度合いに依存するため、分布間の差が不十分に特定される可能性があるという限界を是正すること。
  • カーネルに基づく分布代表(平均埋め込みや滑らかな特徴関数)における $\ell_1$ 幾何を活用して、より強力でスケーラブルな二標本検定を開発すること。
  • 分布の差が生じる場所を明確にする解釈可能なテスト位置を提供することで、モデルの解釈性を向上させること。
  • 最新の $\ell_2$-ベースおよび2次時間計算量のカーネル手法と比較して、統計的検出力と計算効率の両面で優れた性能を示すこと。

提案手法

  • 弱収束をメトリクス化するため、カーネルに基づく分布代表(平均埋め込みまたは滑らかな特徴関数)間の $L^p$ 距離($p \geq 1$)をメトリクスとして用いる。
  • $J$ 個の位置における解析関数の期待値の差を捉えるベクトルの $\ell_1$ ノルムとして、$L^1$ 距離の有限次元近似を提案する。
  • 差がほとんど everywhere で非ゼロとなるように、解析的カーネルを用いて分布代表間の差の稠密性を保証し、$\ell_1$ による検出感度を向上させる。
  • テスト位置を $\ell_1$-ベースの検定統計量を最大化するように最適化することで、解釈可能で識別力のある特徴選択を実現する。
  • ランダムフーリエ特徴を用いて周波数領域に適応させ、線形時間計算を可能にする。
  • サブサンプリングとランダム特徴近似を用いて、一貫性と検出力を維持したままスケーラビリティを確保する。

実験結果

リサーチクエスチョン

  • RQ1$\ell_1$ ノルムの特徴差を用いることで、$\ell_2$-ベースの統計量と比較して、二標本検定における検出力が向上するか?
  • RQ2解析的カーネルの下で差がほとんど everywhere で非ゼロとなるため、$\ell_1$-ベースの検定は分布の差をより効果的に検出できるか?
  • RQ3最新の $\ell_2$-ベースおよび2次時間計算量のカーネル手法と比較して、$\ell_1$-ベースの検定は検出力と速度の両面で優れているか?
  • RQ4選択されたテスト位置は、分布の差が生じる場所を意味的に解釈可能に示せるか?

主な発見

  • 合成データおよび実世界のデータ(20 newsgroups テキストデータセットを含む)において、$\ell_1$-ベースの検定は最新の $\ell_2$-ベース手法よりも高い統計的検出力を達成した。
  • 20 newsgroups データセットにおいて、$\ell_1$-最適化された平均埋め込み検定は、'sci' と 'comp' を区別する際の第2種の誤り確率が 0.00 にまで低下し、'sci' 対 'alt' の場合でも 0.064 に留まった。これは $\ell_2$ 手法を上回る性能を示した。
  • ファストフードレストランの分布タスクでは、$\ell_1$-最適化された検定は名目水準 $\alpha = 0.01$ の近くで第1種の誤り率を維持したが、他の手法はより慎重(保守的)であった。
  • 学習されたテスト位置の可視化から、$\ell_1$-最適化された特徴は分布の重複が少ない領域に集中していることが確認され、その識別力が裏付けられた。
  • $\ell_1$-ベースの手法は、テキストや空間データを含む多様なデータタイプにおいて一貫した性能を示し、2次時間計算量の MMD よりも顕著な高速化を達成した。
  • 非凸最適化の形状により、複数モードで情報が多くなるテスト位置の配置を捉えることができ、凸代替手法に比べてより豊かな検出能力を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。