Skip to main content
QUICK REVIEW

[論文レビュー] Informational Space of Meaning for Scientific Texts

Neslihan Süzen, Evgeny M. Mirkes|arXiv (Cornell University)|Apr 28, 2020
Advanced Text Analysis Techniques参考文献 8被引用数 4
ひとこと要約

本稿では、252のWeb of Science分野にわたる相対的情報利得(RIG)を用いて科学的テキストにおける語の意味を定量化する新しいベクトル空間モデル「意味空間(Meaning Space)」を提案する。レスター科学コーパス(167万件の要約)および辞書(LScDC)に適用したRIGベースの表現は、頻度ベースの手法を上回り、分野特異的かつ高インパクトな科学用語を特定するのに優れている。103,998 × 252のRIG行列と、新たに作成された科学的同義語辞書(LScT)は公開されている。

ABSTRACT

In Natural Language Processing, automatic extracting the meaning of texts constitutes an important problem. Our focus is the computational analysis of meaning of short scientific texts (abstracts or brief reports). In this paper, a vector space model is developed for quantifying the meaning of words and texts. We introduce the Meaning Space, in which the meaning of a word is represented by a vector of Relative Information Gain (RIG) about the subject categories that the text belongs to, which can be obtained from observing the word in the text. This new approach is applied to construct the Meaning Space based on Leicester Scientific Corpus (LSC) and Leicester Scientific Dictionary-Core (LScDC). The LSC is a scientific corpus of 1,673,350 abstracts and the LScDC is a scientific dictionary which words are extracted from the LSC. Each text in the LSC belongs to at least one of 252 subject categories of Web of Science (WoS). These categories are used in construction of vectors of information gains. The Meaning Space is described and statistically analysed for the LSC with the LScDC. The usefulness of the proposed representation model is evaluated through top-ranked words in each category. The most informative n words are ordered. We demonstrated that RIG-based word ranking is much more useful than ranking based on raw word frequency in determining the science-specific meaning and importance of a word. The proposed model based on RIG is shown to have ability to stand out topic-specific words in categories. The most informative words are presented for 252 categories. The new scientific dictionary and the 103,998 x 252 Word-Category RIG Matrix are available online. Analysis of the Meaning Space provides us with a tool to further explore quantifying the meaning of a text using more complex and context-dependent meaning models that use co-occurrence of words and their combinations.

研究の動機と目的

  • 要約のような短い科学的テキストにおける語の意味を定量化する計算モデルの開発。
  • 単なる頻度カウントを超えて、科学文献における分野特異的で意味的に重要な語を同定する課題への対処。
  • 語の意味が分野カテゴリごとの相対的情報利得(RIG)のベクトルとして表現される「意味空間」の構築。
  • 自然言語処理およびテキストマイニング応用に利用可能な、公開可能な科学的辞書とRIG行列の作成。

提案手法

  • 各語を252のWeb of Science分野カテゴリにわたる相対的情報利得(RIG)値のベクトルとして表現する。
  • 情報理論的原則を用いて、語の存在がカテゴリに関する不確実性をどの程度低減するかを測定することで、各語-カテゴリペアのRIGを計算する。
  • 1,673,350件の科学的要約から構成されるレスター科学コーパス(LSC)と、それをもとに作成されたレスター科学辞書コア(LScDC)を用いて意味空間を構築する。
  • RIG行列を用いて各カテゴリごとに情報量の多い語をランク付けし、分野特異的用語の同定を可能にする。
  • 次元削減およびクラスタリング技術を用いて、意味空間の構造を探索する。
  • 103,998 × 252の語-カテゴリRIG行列と、新たに作成された科学的同義語辞書(LScT)を公開する。

実験結果

リサーチクエスチョン

  • RQ1RIGベースの語表現は、科学的テキストにおける意味的・分野特異的語を特定する際、頻度ベースの手法を上回るか?
  • RQ2RIGベースのベクトル空間モデルは、多様な研究分野にわたる科学的用語の意味的特異性をどの程度適切に捉えられるか?
  • RQ3RIGを用いて構築された意味空間は、語の使用法およびカテゴリとの関連性に有意義なパターンを明らかにするか?
  • RQ4RIG行列およびそれによって得られる同義語辞書(LScT)は、科学的テキスト解析および情報抽出のための堅牢でデータ駆動型のツールとして機能するか?

主な発見

  • RIGベースのランク付けは、252のすべての分野カテゴリにおいて、分野特異的かつ高インパクトな科学用語を特定する際、頻度ベースの手法を顕著に上回った。
  • 各カテゴリで最も情報量の多い語—例えば、女性学分野の「femal」(RIG: 3.6×10⁻²)や動物学分野の「speci」(RIG: 1.9×10⁻¹)—は、非常に関連性が高く文脈的に意味のある語であった。
  • 103,998 × 252の語-カテゴリRIG行列は、分野を越えた科学的語の意味を包括的かつ公開可能な形で表現している。
  • 提案された意味空間は、情報利得分析を用いて異常や分野特異的用語の発見に効果的に機能する。
  • RIG行列から、レスター科学的同義語辞書(LScT)が成功裏に生成され、科学的テキストマイニングのための新たなリソースが得られた。
  • 本モデルは、語の共起や意味的組み合わせを含む高度な意味モデルへの応用の可能性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。