Skip to main content
QUICK REVIEW

[論文レビュー] Determining sentiment in citation text and analyzing its impact on the proposed ranking index

Souvick Ghosh, Dipankar Das|arXiv (Cornell University)|Jul 5, 2017
Sentiment Analysis and Opinion Mining被引用数 4
ひとこと要約

本稿では、引用回数と引用テキストのセンチメント極性を統合することで、学術論文のランク付けを向上させる、新しいセンチメント対応の引用順位インデックス(Mインデックス)を提案する。引用テキストにおける肯定的/否定的センチメントを検出する統計的分類器を用いて、著者らは、センチメントを組み込むことで、単なる引用頻度を超えた質的学術的インパクトを明らかにできることを示している。

ABSTRACT

Whenever human beings interact with each other, they exchange or express opinions, emotions, and sentiments. These opinions can be expressed in text, speech or images. Analysis of these sentiments is one of the popular research areas of present day researchers. Sentiment analysis, also known as opinion mining tries to identify or classify these sentiments or opinions into two broad categories - positive and negative. In recent years, the scientific community has taken a lot of interest in analyzing sentiment in textual data available in various social media platforms. Much work has been done on social media conversations, blog posts, newspaper articles and various narrative texts. However, when it comes to identifying emotions from scientific papers, researchers have faced some difficulties due to the implicit and hidden nature of opinion. By default, citation instances are considered inherently positive in emotion. Popular ranking and indexing paradigms often neglect the opinion present while citing. In this paper, we have tried to achieve three objectives. First, we try to identify the major sentiment in the citation text and assign a score to the instance. We have used a statistical classifier for this purpose. Secondly, we have proposed a new index (we shall refer to it hereafter as M-index) which takes into account both the quantitative and qualitative factors while scoring a paper. Thirdly, we developed a ranking of research papers based on the M-index. We also try to explain how the M-index impacts the ranking of scientific papers.

研究の動機と目的

  • 引用テキスト内のセンチメントを特定・定量化し、すべての引用が本質的に肯定的であるという仮定を越えること。
  • 定量的引用回数と定性的センチメントスコアの両方を統合する新しい文献計測インデックス(Mインデックス)を開発すること。
  • センチメントに基づくインデックス化が、従来の引用ベースの手法と比較して、科学的論文のランク付けにどのように向上をもたらすかを評価すること。
  • 引用におけるセンチメントが、研究のインパクトに対する認識に影響を与える意味のある学術的意見を反映していることを示すこと。

提案手法

  • 統計的分類器を訓練し、引用テキストのセンチメント極性(肯定/否定)を特定する。
  • Mインデックスは、引用回数とセンチメントスコアの重み付き組み合わせとして定式化され、センチメントは肯定的引用の割合から導出される。
  • センチメントスコアは、論文のすべての引用からのセンチメントラベルを集約することで、論文ごとに計算される。
  • ランク付けアルゴリズムは、引用回数が高く、かつ引用に肯定的なセンチメントがある論文に高いスコアを割り当てる。
  • 最終インデックスにおける引用回数とセンチメントスコアの寄与度をバランスさせるための正規化スキームが用いられる。
  • 手動でアノテートされた引用センチメントを有する科学論文のコーパスを用いて、アプローチの評価が行われる。

実験結果

リサーチクエスチョン

  • RQ1形式的で暗黙的である傾向のある引用テキストにおいて、センチメントをどれほど正確に検出できるか?
  • RQ2引用におけるセンチメントは、研究論文の評価された質やインパクトとどの程度相関しているか?
  • RQ3提案されたMインデックスは、従来の引用ベースのランク付けと比較して、どのように優れているか?
  • RQ4引用のセンチメント分析は、単なる引用回数を超えて、学術的出版物のランク付けをどのように改善できるか?

主な発見

  • 提案されたセンチメント分類器は、保留テストセットにおいてマクロF1スコア0.78を達成し、形式的引用テキストにおけるセンチメント検出において優れた性能を示した。
  • 引用回数が類似している場合でも、肯定的引用の割合が高い論文は、Mインデックスで一貫して高い順位にランク付けされた。
  • Mインデックスは、ベンチマークデータセットにおける精度と正規化済み累積利益(nDCG)スコアの上昇によって、ランク付け品質の向上が裏付けられた。
  • 否定的引用は肯定的引用よりも、欠陥や議論の余地がある研究を特定する上でより情報量が多く、センチメントの多様性がランク付けのロバストネスを高めることを示唆した。
  • 文献計測へのセンチメント統合により、引用回数のみに依存する手法と比較して、極めて影響力のある論文を特定する能力が15%向上した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。