Skip to main content
QUICK REVIEW

[論文レビュー] Nonparametric Detection of Anomalous Data via Kernel Mean Embedding.

Shaofeng Zou, Yingbin Liang|arXiv (Cornell University)|Apr 25, 2014
Statistical Methods and Inference参考文献 21被引用数 8
ひとこと要約

本稿では、再生核ヒルバート空間(RKHS)におけるカーネル平均埋め込みを用いた非パrametricな異常検出手法を提案する。最大平均差分(MMD)を用いて、n個の系列の中からs個の異常系列を検出する。sが既知の場合、各系列にm = O(log n)のサンプルがあれば一貫した検出が可能であり、sが未知の場合にはm > O(log n)が必要である。計算量は多項式時間であり、従来の手法およびカーネルベースの手法と比較して優れた実験的性能を示す。

ABSTRACT

An anomaly detection problem is investigated, in which there are totally n sequences with s anomalous sequences to be detected. Each normal sequence contains m independent and identically distributed (i.i.d.) samples drawn from a distribution p, whereas each anomalous sequence contains m i.i.d. samples drawn from a distribution q that is distinct from p. The distributions p and q are assumed to be unknown a priori. Two scenarios, respectively with and without a reference sequence generated by p, are studied. Distribution-free tests are constructed using maximum mean discrepancy (MMD) as the metric, which is based on mean embeddings of distributions into a reproducing kernel Hilbert space (RKHS). For both scenarios, it is shown that as the number n of sequences goes to infinity, if the value of s is known, then the number m of samples in each sequence should be at the order O(log n) or larger in order for the developed tests to consistently detect s anomalous sequences. If the value of s is unknown, then m should be at the order strictly larger than O(log n). Computational complexity of all developed tests is shown to be polynomial. Numerical results demonstrate that our tests outperform (or perform as well as) the tests based on other competitive traditional statistical approaches and kernel-based approaches under various cases. Consistency of the proposed test is also demonstrated on a real data set.

研究の動機と目的

  • 基礎的な分布pとqが未知である状況において、n個の系列の中からs個の異常系列を検出する課題に取り組む。
  • 最大平均差分(MMD)に基づく分布フリーな仮説検定を、pからの参照系列がある場合とない場合の両方の状況で開発する。
  • 既知および未知のsにおける一貫した検出のための理論的サンプルサイズ要件を確立する。
  • すべての提案テストにおいて多項式時間の計算複雑度を維持することで、計算効率を確保する。
  • 合成データおよび実データ上で、従来の統計的手法およびカーネルベースの手法と比較して、本手法の性能を検証する。

提案手法

  • 再生核ヒルバート空間(RKHS)におけるカーネル平均埋め込みを用いて、確率分布pとqをその空間の要素として表現し、非パラメトリックな比較を可能にする。
  • 正規系列と異常系列の平均埋め込み間の統計的距離として、最大平均差分(MMD)を用いる。
  • 正規分布pからの逸脱を検出するための分布フリーな仮説検定をMMDに基づいて構築する。
  • 参照系列が存在する場合とない場合の2種類のテストバリエーションを設計し、それに応じて検定統計量を調整する。
  • n → ∞の下で検出の一貫性を保証するための、各系列あたりのサンプル数mに関する理論的条件を導出する。
  • 計算複雑度を分析し、すべての提案テストにおいてnおよびmに関して多項式時間の複雑度を維持することを証明する。

実験結果

リサーチクエスチョン

  • RQ1sが既知である場合、未知の分布pとqのもとで、s個の異常系列を一貫して検出するために各系列に必要な最小サンプル数mはどれほどか?
  • RQ2sが未知の場合、必要なサンプル数mはどのように変化するのか。また、どのような理論的保証を提供できるか?
  • RQ3MMDに基づくテストは、pとqのパラメトリックな形を仮定せずに一貫した異常検出を達成できるか?
  • RQ4本手法は、従来の手法およびカーネルベースの異常検出手法と比較して、性能および複雑度の面でどのように差をつけるか?
  • RQ5未知の基礎的分布を有する実世界のデータに対しても、本手法は一貫性と有効性を維持するか?

主な発見

  • sが既知の場合、各系列にm = O(log n)のサンプルがあれば、n → ∞の下でs個の異常系列の一貫した検出が可能である。
  • sが未知の場合、一貫した検出を保証するためにはmがO(log n)より厳密に大きくなければならない。
  • すべての提案テストは多項式時間の計算複雑度を有しており、nが大きい場合でもスケーラブルである。
  • 数値実験の結果、提案されたMMDベースのテストは、さまざまな設定において従来の統計的手法およびカーネルベースの手法を上回るか、同等の性能を示した。
  • 実世界のデータセットにおいても本手法は一貫性を示し、実用的妥当性および理論的整合性を確認した。
  • 理論的枠組みは、参照系列がある場合とない場合の両方をうまく処理でき、分布の不確実性下でも頑健な検出を提供した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。