Skip to main content
QUICK REVIEW

[論文レビュー] Clustering With Side Information: From a Probabilistic Model to a Deterministic Algorithm

Daniel Khashabi, John Wieting|arXiv (Cornell University)|Aug 25, 2015
Bayesian Methods and Mixture Models参考文献 24被引用数 11
ひとこと要約

本稿では、ノイズの多い状況下でもクラスタリングのロバスト性を向上させるために、データと補助情報(対比較制約)を同時にモデル化する非パrametric Bayesianモデル、TVClustを提案する。このモデルから、クラスタ数を事前に指定しなくてもよい関係的補助情報を組み込んだ決定的クラスタリング手法RDP-meansを導出しており、従来の制約付きクラスタリング手法と比較して、ノイズや誤った入力に対し優れた性能を発揮する。

ABSTRACT

In this paper, we propose a model-based clustering method (TVClust) that robustly incorporates noisy side information as soft-constraints and aims to seek a consensus between side information and the observed data. Our method is based on a nonparametric Bayesian hierarchical model that combines the probabilistic model for the data instance and the one for the side-information. An efficient Gibbs sampling algorithm is proposed for posterior inference. Using the small-variance asymptotics of our probabilistic model, we then derive a new deterministic clustering algorithm (RDP-means). It can be viewed as an extension of K-means that allows for the inclusion of side information and has the additional property that the number of clusters does not need to be specified a priori. Empirical studies have been carried out to compare our work with many constrained clustering algorithms from the literature on both a variety of data sets and under a variety of conditions such as using noisy side information and erroneous k values. The results of our experiments show strong results for our probabilistic and deterministic approaches under these conditions when compared to other algorithms in the literature.

研究の動機と目的

  • ノイズが多く、ヒューリスティックに基づく補助情報をクラスタリングに組み込む際、それを真のラベルとして扱わない課題に対処すること。
  • データと対比較制約の両方から同時に学習する原理的で確率的モデルを構築すること。
  • 確率的モデルの漸近的解析から、クラスタ数の自動推定が可能な決定的クラスタリングアルゴリズムを導出すること。
  • ノイズの多い補助情報や誤ったk値に対するロバスト性を評価すること。これは、現実世界のクラスタリングにおいて一般的な問題である。

提案手法

  • データのためのディリクレ過程混合モデルと補助情報のためのランダムグラフモデルを組み合わせた二視点非パrametric Bayesianモデル、TVClustを提案する。
  • TVClustモデルにおける事後分布推論にギブスサンプリングを用い、クラスタ割り当てとモデルパラメータを推定する。
  • TVClustモデルに小分散漸近解析を適用し、関係的制約を組み込んだDP-meansの一般化である決定的アルゴリズムRDP-meansを導出する。
  • RDP-meansは、多項分布データ(例:SIFT特徴量)に対してKLダイバージェンスを用い、必須リンク制約を関係性ペナルティ項として統合する。
  • アルゴリズムは、近接性と制約の整合性に基づいてコンポーネントを統合することで、クラスタ数を動的に決定する。
  • 離散的特徴空間におけるKLダイバージェンス計算の安定化のため、ラプラススムージング(α=0.3)を採用する。

実験結果

リサーチクエスチョン

  • RQ1補助情報がノイズが多く、完全に信頼できない状況下で、どのようにしてそれをクラスタリングにロバストに組み込むことができるか?
  • RQ2データと制約を同時にモデル化する原理的な確率的モデルは、既存の制約付きクラスタリング手法を上回る性能を発揮できるか?
  • RQ3ベイジアンモデルの小分散漸近近似から決定的アルゴリズムを導出することで、誤ったk値に対しても効果的かつロバストな手法が得られるか?
  • RQ4制約数が限られている状況下で、関係的補助情報を組み込むとクラスタリング性能にどのように影響するか?

主な発見

  • UCIデータセットにおいて、高ノイズ(p=1, r=0.03)条件下でRDP-meansはF-measure 0.96を達成し、同条件でMPCKMeans(0.58)とLCVQE(0.86)を著しく上回った。
  • ImageNetデータセットでは、多項分布モデルを用いたRDP-meansがF-measure 0.44を達成し、ガウスモデル(0.20)とK-means(0.25)を上回った。
  • RDP-meansはk値の変動に対しても高い安定性を示し、kが±3ずれてもF-measureが0.82以上を維持しており、誤ったk値入力に対するロバスト性を実証した。
  • RDP-meansの性能は、制約のサンプリングレートが上昇するにつれて単調に向上し、全ペアの6%程度のサンプリングでもほぼ最適な結果に到達した。
  • TVClust(変分推論)はUCIデータセット全体でF-measure 0.72~0.79を達成し、ノイズの多い制約下でも強力な性能を示した。
  • 本手法は、ノイズの多い補助情報と誤ったk値の両方に対してロバストであり、RDP-meansとTVClustともに摂動に対して最小限の性能低下を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。