Skip to main content
QUICK REVIEW

[論文レビュー] A Unified Subspace Outlier Ensemble Framework for Outlier Detection in High Dimensional Spaces

Zengyou He, Xiaofei Xu|ArXiv.org|May 24, 2005
Anomaly Detection Techniques and Applications被引用数 3
ひとこと要約

本稿では、高次元空間における外れ値検出のための統一的サブスペース外れ値アンサンブルフレームワークを提案する。外れ値スコアは、アンサンブル学習を用いてサブスペース固有の要因の統合としてモデル化される。1次元サブスペースのみを用いるSOE1アルゴリズムは、大規模なカテゴリカルデータセットにおいて、既存の手法と比較して最大10倍速く、最先端の外れ値検出性能を達成する。

ABSTRACT

The task of outlier detection is to find small groups of data objects that are exceptional when compared with rest large amount of data. Detection of such outliers is important for many applications such as fraud detection and customer migration. Most such applications are high dimensional domains in which the data may contain hundreds of dimensions. However, the outlier detection problem itself is not well defined and none of the existing definitions are widely accepted, especially in high dimensional space. In this paper, our first contribution is to propose a unified framework for outlier detection in high dimensional spaces from an ensemble-learning viewpoint. In our new framework, the outlying-ness of each data object is measured by fusing outlier factors in different subspaces using a combination function. Accordingly, we show that all existing researches on outlier detection can be regarded as special cases in the unified framework with respect to the set of subspaces considered and the type of combination function used. In addition, to demonstrate the usefulness of the ensemble-learning based outlier detection framework, we developed a very simple and fast algorithm, namely SOE1 (Subspace Outlier Ensemble using 1-dimensional Subspaces) in which only subspaces with one dimension is used for mining outliers from large categorical datasets. The SOE1 algorithm needs only two scans over the dataset and hence is very appealing in real data mining applications. Experimental results on real datasets and large synthetic datasets show that: (1) SOE1 has comparable performance with respect to those state-of-art outlier detection algorithms on identifying true outliers and (2) SOE1 can be an order of magnitude faster than one of the fastest outlier detection algorithms known so far.

研究の動機と目的

  • 高次元空間における外れ値検出のための広く受け入れられた定義の欠如に対処すること。
  • 既存の外れ値検出手法を1つのアンサンブル学習フレームワークに統合すること。
  • 最小限のサブスペース次元数を用いて、大規模なカテゴリカルデータセットのための高速でスケーラブルなアルゴリズムを開発すること。
  • サブスペースベースのアンサンブル手法が、顕著に低い計算コストで高い精度を達成できることを示すこと。

提案手法

  • フレームワークは、複数のサブスペースにわたる外れ値要因の組み合わせとして外れ値スコアを定義する。
  • 異なるサブスペースからの外れ値スコアを統合する関数を採用し、既存の手法の統一的取り扱いを可能にする。
  • SOE1アルゴリズムはサブスペースを1次元に制限することで、計算を単純化し、I/Oオーバーヘッドを低減する。
  • SOE1はデータセットに対してたった2回のフルスキャンしか行わず、大規模なデータマイニングにおいて極めて効率的である。
  • 外れ値検出は、1次元サブスペースにおける頻度に基づく外れ値要因に基づき、統合関数を介して集約される。
  • サブスペースの集合と統合関数を変更することで、既存の手法が特別なケースとして一般化される。

実験結果

リサーチクエスチョン

  • RQ1高次元空間における外れ値検出を、統一的フレームワークの下で形式化することは可能か?
  • RQ2サブスペースにおけるアンサンブル学習は、検出精度と効率性を向上させることができるか?
  • RQ31次元サブスペースのみを用いる場合の、精度と速度のトレードオフはいかなるものか?
  • RQ4提案されたフレームワークは、既存の外れ値検出アルゴリズムとどのように関係し、一般化されるか?

主な発見

  • SOE1は、実データおよび合成データセットにおいて、最先端のアルゴリズムと同等の外れ値検出性能を達成する。
  • SOE1は、既存の高速な外れ値検出アルゴリズムの中でも最も速いものの1つと比較して、最大10倍速い。
  • 統一されたフレームワークは、サブスペース選択と統合関数の変更により、すべての既存の外れ値検出手法を特別なケースとして包含する。
  • SOE1における1次元サブスペースの使用により、大規模なデータマイニングに適した極めて効率的な2スキャンアルゴリズムが実現される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。