Skip to main content
QUICK REVIEW

[論文レビュー] MMD Aggregated Two-Sample Test

Antonin Schrab, Ilmun Kim|arXiv (Cornell University)|Oct 28, 2021
Radiomics and Machine Learning in Medical Imaging参考文献 48被引用数 11
ひとこと要約

本稿では、最大平均差分(MMD)に基づく新しい非パラメトリックな2標本検定、MMDAggを提案する。この手法は、保持データやヒューリスティックな選択を必要とせず、複数のカーネル帯域幅を適応的に集約する。非漸近的な第一種誤り率の制御と、反復対数項を除けば最小最大最適性を達成し、合成データおよび実世界の画像データにおいて、最先端のMMDベースの検定を上回る性能を示す。

ABSTRACT

We propose two novel nonparametric two-sample kernel tests based on the Maximum Mean Discrepancy (MMD). First, for a fixed kernel, we construct an MMD test using either permutations or a wild bootstrap, two popular numerical procedures to determine the test threshold. We prove that this test controls the probability of type I error non-asymptotically. Hence, it can be used reliably even in settings with small sample sizes as it remains well-calibrated, which differs from previous MMD tests which only guarantee correct test level asymptotically. When the difference in densities lies in a Sobolev ball, we prove minimax optimality of our MMD test with a specific kernel depending on the smoothness parameter of the Sobolev ball. In practice, this parameter is unknown and, hence, the optimal MMD test with this particular kernel cannot be used. To overcome this issue, we construct an aggregated test, called MMDAgg, which is adaptive to the smoothness parameter. The test power is maximised over the collection of kernels used, without requiring held-out data for kernel selection (which results in a loss of test power), or arbitrary kernel choices such as the median heuristic. We prove that MMDAgg still controls the level non-asymptotically, and achieves the minimax rate over Sobolev balls, up to an iterated logarithmic term. Our guarantees are not restricted to a specific type of kernel, but hold for any product of one-dimensional translation invariant characteristic kernels. We provide a user-friendly parameter-free implementation of MMDAgg using an adaptive collection of bandwidths. We demonstrate that MMDAgg significantly outperforms alternative state-of-the-art MMD-based two-sample tests on synthetic data satisfying the Sobolev smoothness assumption, and that, on real-world image data, MMDAgg closely matches the power of tests leveraging the use of models such as neural networks.

研究の動機と目的

  • 既存のMMDベースの2標本検定において、特に小標本設定下で非漸近的な第一種誤り率の制御が欠如している問題に対処すること。
  • 未知の滑らかさパラメータを伴うSobolev球上でのカーネル選択に、保持データを必要とせずに適応する検定を開発すること。
  • 有限標本の妥当性とロバスト性を維持しながら、Sobolev球上での最小最大最適性を達成すること。
  • 自動的な帯域幅集合を用いた、適応的MMD検定の実用的で使いやすい実装を提供すること。

提案手法

  • パーミュテーションまたはワイルドブートストラップを用いて閾値をキャリブレーションする非漸近的MMD検定を提案し、有限標本サイズ下での正確な第一種誤り率制御を保証する。
  • 既知の滑らかさパラメータ s に対して、最適帯域幅を用いた固定帯域幅の単一MMD検定を導入し、Sobolev球上での最小最大最適性を達成する。
  • データ駆動型の帯域幅集合を用いて、MMD統計量を適応的重みで集約することでMMDAggを構築する。この重みは検定力の最大化を目的とする。
  • バイアスと分散のバランスを最適化するために選ばれる、dyadicスケーリングに基づくパラメータフリーの帯域幅集合(λ ∈ {2^{-ℓ} : ℓ = 0,…,ℓ*})を採用する。
  • 帰無仮説下での期待分散の逆数から導かれる重みを用いたMMD統計量の重み付き組み合わせにより、ロバストな集約を実現する。
  • MMDAggが非漸近的に第一種誤り率を制御し、Sobolev球上での最小最大分離レートを反復対数因子を除いて達成することを証明する。

実験結果

リサーチクエスチョン

  • RQ1MMDに基づく非パラメトリック2標本検定は、有限標本サイズ下でも正確な第一種誤り率制御を維持できるか?
  • RQ2滑らかさパラメータ s を事前に知らない状況でも、Sobolev球上での最小最大最適性を達成できる適応的MMD検定を構築できるか?
  • RQ3カーネル選択のための保持データによるパワー損失や、メジアンヒューリスティックのような任意の選択を回避できるか?
  • RQ4集約されたMMD検定は、非漸近的な水準制御を維持しながら、近似的最適な検出率を達成できるか?
  • RQ5実世界の画像データにおいて、MMDAggはニューラルネットワークのようなモデルベース手法と比較してどのように性能を発揮するか?

主な発見

  • 提案されたMMDAgg検定は非漸近的に第一種誤り率を制御し、小標本サイズ下でも信頼性の高いキャリブレーションを実現する。
  • 既知の滑らかさパラメータ s に対して、最適帯域幅を用いた単一MMD検定はSobolev球上での最小最大最適性を達成する。
  • MMDAggはSobolev球上での最小最大分離レートを、(ln ln(m+n))^{-2s/(4s+d)} の要因を除いて達成する。これは対数項を除いて最適である。
  • Sobolev球のクラス {S_d^s(R) : s > 0, R > 0} 上で最小最大適応的であるが、s や R の知識を必要としない。
  • Sobolev滑らかさ下での合成データにおいて、MMDAggは既存のMMDベースの検定を著しく上回るパワーを示す。
  • MNISTデータセットでは、MMDAggは深層学習ベースの2標本検定と同等のパワーを達成し、実世界の画像シフト検出において優れた経験的性能を示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。