Skip to main content
QUICK REVIEW

[論文レビュー] A Semantic approach for effective document clustering using WordNet

Leena H. Patil, Mohammed Atique|arXiv (Cornell University)|Mar 3, 2013
Text and Document Classification Technologies参考文献 6被引用数 5
ひとこと要約

この論文では、語彙的関係と次元削減を向上させるために WordNet を用いた意味的文書クラスタリング手法を提案する。WordNet を用いた類義語関係および意味的関係の統合、前処理(ストップワード除去、ステミング)の適用、および tf-idf、tf-df、tf2 を用いた語の重み付けにより、Reuters-21578 や 20 Newsgroups といったベンチマークデータセットにおけるクラスタリング精度が向上する。

ABSTRACT

Now a days, the text document is spontaneously increasing over the internet, e-mail and web pages and they are stored in the electronic database format. To arrange and browse the document it becomes difficult. To overcome such problem the document preprocessing, term selection, attribute reduction and maintaining the relationship between the important terms using background knowledge, WordNet, becomes an important parameters in data mining. In these paper the different stages are formed, firstly the document preprocessing is done by removing stop words, stemming is performed using porter stemmer algorithm, word net thesaurus is applied for maintaining relationship between the important terms, global unique words, and frequent word sets get generated, Secondly, data matrix is formed, and thirdly terms are extracted from the documents by using term selection approaches tf-idf, tf-df, and tf2 based on their minimum threshold value. Further each and every document terms gets preprocessed, where the frequency of each term within the document is counted for representation. The purpose of this approach is to reduce the attributes and find the effective term selection method using WordNet for better clustering accuracy. Experiments are evaluated on Reuters Transcription Subsets, wheat, trade, money grain, and ship, Reuters 21578, Classic 30, 20 News group (atheism), 20 News group (Hardware), 20 News group (Computer Graphics) etc.

研究の動機と目的

  • ウェブやメールシステムのような大規模なテキストリポジトリにおける効果的な文書クラスタリングの課題に対処すること。
  • 外部知識としての WordNet を用いて語の意味的関係を活用することで、クラスタリング精度を向上させること。
  • 意味的および統計的基準に基づく知的な語の選択により、特徴空間の次元を削減すること。
  • 伝統的な語の重み付け手法(tf-idf、tf-df、tf2)と WordNet を組み合わせた場合の文書クラスタリングにおける有効性を評価すること。
  • 意味的認識のある前処理および特徴選択を用いることで、標準的なテキストクラスタリングベンチマークで優れたパフォーランスを示すこと。

提案手法

  • 文書をストップワード除去およびポーターステミングを用いて正規化することで前処理を行う。
  • WordNet を活用して類義語および意味的関係を特定し、グローバルに一意の語と頻出語の集合を生成する。
  • 前処理済みおよび意味的に豊かにされた語を用いて、語-文書行列を構築する。
  • 複数の語の選択手法(tf-idf、tf-df、tf2)を用い、最小閾値を設定して関連語を選択する。
  • 各文書内の語の頻度を数え、データ行列への表現を実現する。
  • WordNet からの意味的知識を統合して語の表現を強化し、クラスタリング性能を向上させる。

実験結果

リサーチクエスチョン

  • RQ1WordNet の意味的関係を組み込むことで、従来手法と比較して文書クラスタリング精度がどの程度向上するか?
  • RQ2意味的拡張と組み合わせた場合、語の選択戦略(tf-idf、tf-df、tf2)の中でどの手法が最も優れたクラスタリングパフォーマンスを示すか?
  • RQ3意味的前処理は、関連する文書特徴を保持しつつ、どの程度次元削減を実現できるか?
  • RQ4Reuters-21578 や 20 Newsgroups のような多様なテキストコレクションにおいて、提案手法はどの程度のパフォーマンスを示すか?
  • RQ5WordNet からの意味的関係は、クラスタリングタスクにおける語の表現を効果的に向上させることができるか?

主な発見

  • WordNet の統合により、意味的関係を用いた語の表現が豊かにされ、クラスタリング精度が顕著に向上した。
  • 意味的前処理と組み合わせた場合、tf-df および tf2 手法が tf-idf よりもクラスタリング品質において優れていた。
  • 提案手法は、選択された語の高い関連性を維持しつつ、特徴空間の次元を効果的に削減した。
  • Reuters-21578 および 20 Newsgroups データセットにおける実験により、F-measure および純度の両面で一貫した向上が確認された。
  • WordNet を用いた意味的語の拡張により、グローバル語および頻出語の集合の特定がより良く行われ、全体のクラスタリングパフォーマンスが向上した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。