Skip to main content
QUICK REVIEW

[論文レビュー] A Novel Information Theoretic Framework for Finding Semantic Similarity in WordNet

Abhijit Adhikari, S. Paul Singh|arXiv (Cornell University)|Jul 19, 2016
Topic Modeling参考文献 1被引用数 7
ひとこと要約

本稿は、コーパスに依存しない内在的情報含量(IC)計算モデルを用いて、WordNetにおける意味的類似度のための情報理論的フレームワークを提案する。WordNetのオントロジーのトポロジカルな性質を活用することで、ゴールスタンダードデータセットとの相関が高く(最大0.71)得られ、最先端のICベースおよび非ICベースの手法を上回る性能を発揮する。

ABSTRACT

Information content (IC) based measures for finding semantic similarity is gaining preferences day by day. Semantics of concepts can be highly characterized by information theory. The conventional way for calculating IC is based on the probability of appearance of concepts in corpora. Due to data sparseness and corpora dependency issues of those conventional approaches, a new corpora independent intrinsic IC calculation measure has evolved. In this paper, we mainly focus on such intrinsic IC model and several topological aspects of the underlying ontology. Accuracy of intrinsic IC calculation and semantic similarity measure rely on these aspects deeply. Based on these analysis we propose an information theoretic framework which comprises an intrinsic IC calculator and a semantic similarity model. Our approach is compared with state of the art semantic similarity measures based on corpora dependent IC calculation as well as intrinsic IC based methods using several benchmark data set. We also compare our model with the related Edge based, Feature based and Distributional approaches. Experimental results show that our intrinsic IC model gives high correlation value when applied to different semantic similarity models. Our proposed semantic similarity model also achieves significant results when embedded with some state of the art IC models including ours.

研究の動機と目的

  • 従来の意味的類似度のための情報含量(IC)計算におけるデータスパarsityおよびコーパス依存性の問題に対処すること。
  • WordNetのトポロジカル構造に基づく、コーパスに依存しない内在的ICモデルの開発。
  • 内在的ICとオントロジーのトポロジーを統合することで、意味的類似度測定の精度を向上させること。
  • 提案されたフレームワークを、最先端のICベースおよび非ICベースの意味的類似度モデルと比較して評価すること。
  • MengらやPirróらの他のICモデルと統合した場合の互換性と優れた性能を示すこと。

提案手法

  • 外部コーパスに依存せずに、WordNetの階層的構造から情報含量を導出する新しい内在的IC計算手法を提案する。
  • ルートからのパス長、祖先の数、深さなどのトポロジカル特徴を用いて、内在的IC値を計算する。
  • 標準的な類似度関数(例:Resnik、Lin、Jiang-Conrath)を用いた意味的類似度フレームワークに、内在的ICモデルを統合する。
  • 人間によるアノテーションがなされた類似度スコアを有するベンチマークデータセットを用いたマルチソース評価戦略を採用する。
  • 複数のIC計算手法(Secoら、Zhouら、Sánchezら、Mengら、Qingboら)および類似度モデルを用いてフレームワークを検証する。
  • 多様なICおよび類似度モデルにおいて、ゴールスタンダードデータとの相関係数を比較するための統一評価パイプラインを適用する。

実験結果

リサーチクエスチョン

  • RQ1単にWordNetのトポロジーに基づく内在的ICモデルは、コーパス依存のIC手法を上回る性能を示せるか?
  • RQ2提案された内在的ICモデルは、既存の意味的類似度測定法(例:Resnik、Lin、Jiang-Conrath)の性能にどのように影響を与えるか?
  • RQ3提案されたフレームワークは、最先端のICベースおよび非ICベースのモデルと比較して、ゴールスタンダードデータセットとどの程度相関を示すか?
  • RQ4Meng らやPirró のような他のICモデルと統合した場合、提案されたフレームワークはどの程度の頑健性を示すか?
  • RQ5内在的ICモデルは、さまざまな類似度関数に一般化可能であり、高い精度を維持できるか?

主な発見

  • 提案された内在的ICモデルは、Pirró類似度モデルと組み合わせた場合、ゴールスタンダードデータと0.71の相関を示し、テストされたすべてのICベースのモデルの中で最高の性能を発揮した。
  • 自らの類似度モデルを用いた場合、ゴールスタンダードデータセットと0.69の相関を示し、強力な性能を示した。
  • Meng らのICモデルと統合した場合、類似度モデルは0.68の相関を達成し、高い互換性と有効性を示した。
  • 内在的ICモデルは、さまざまな類似度関数において、コーパス依存のIC手法と比較して、安定性と一般化性能の面で常に優れていた。
  • フレームワークは高い互換性を示し、既存のICモデルおよび類似度関数と組み合わせた場合、最高の結果を達成した。
  • 提案手法は、評価されたすべての内在的ICベースのモデルの中で最高の相関(0.71)を達成し、意味的類似度推定における優位性を確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。