Skip to main content
QUICK REVIEW

[論文レビュー] Persian Sentiment Analyzer: A Framework based on a Novel Feature Selection Method

Ayoub Bagheri, Mohamad Saraee|arXiv (Cornell University)|Dec 27, 2014
Sentiment Analysis and Opinion Mining参考文献 28被引用数 13
ひとこと要約

本稿では、語形変化、不規則な空白、口語的表現といった課題に対処するため、ペルシャ語のセンチメント分析のための新規特徴選択フレームワークを提案する。語彙還元と独自の特徴選択手法を組み合わせ、ナイーブベイズを分類に用いることで、手動で収集したペルシャ語のスマートフォンレビューのデータセット上で、分類精度が向上し、低リソースNLP環境における本手法の有効性が示された。

ABSTRACT

In the recent decade, with the enormous growth of digital content in internet and databases, sentiment analysis has received more and more attention between information retrieval and natural language processing researchers. Sentiment analysis aims to use automated tools to detect subjective information from reviews. One of the main challenges in sentiment analysis is feature selection. Feature selection is widely used as the first stage of analysis and classification tasks to reduce the dimension of problem, and improve speed by the elimination of irrelevant and redundant features. Up to now as there are few researches conducted on feature selection in sentiment analysis, there are very rare works for Persian sentiment analysis. This paper considers the problem of sentiment classification using different feature selection methods for online customer reviews in Persian language. Three of the challenges of Persian text are using of a wide variety of declensional suffixes, different word spacing and many informal or colloquial words. In this paper we study these challenges by proposing a model for sentiment classification of Persian review documents. The proposed model is based on lemmatization and feature selection and is employed Naive Bayes algorithm for classification. We evaluate the performance of the model on a manually gathered collection of cellphone reviews, where the results show the effectiveness of the proposed approaches.

研究の動機と目的

  • ペルシャ語のセンチメント分析における特徴選択に関する研究の不足に取り組む。
  • 語形変化接尾語、可変的空白、口語的語彙を含むペルシャ語テキストの言語的課題に対処する。
  • ペルシャ語のオンラインレビューに特化したセンチメント分類フレームワークを開発する。
  • 低リソース環境における分類性能に与える、異なる特徴選択手法の影響を評価する。
  • 手動で整備したペルシャ語スマートフォンレビューのデータセット上で、提案モデルの有効性を示す。

提案手法

  • フレームワークは、ペルシャ語の語形を正規化し、語形のばらつきを低減するために語彙還元を適用する。
  • センチメント分類に最も関連する特徴を特定・保持するための新規特徴選択手法を提案する。
  • 不要で重複する特徴を削除することで次元削減を実現し、計算効率を向上させる。
  • テキスト分類タスクにおいて有効であるとして、分類アルゴリズムにナイーブベイズを用いる。
  • モデルは、手動で収集したペルシャ語スマートフォンレビューのデータセット上で学習および評価される。
  • 分類性能は標準的な分類指標を用いて測定され、ベースライン手法と比較して精度が向上していることが示された。

実験結果

リサーチクエスチョン

  • RQ1提案された特徴選択手法は、ペルシャ語テキストのセンチメント分類精度において、既存手法と比較してどのように差がつくか?
  • RQ2語形変化が著しいペルシャ語レビューにおいて、語彙還元はどの程度センチメント分類性能を向上させるか?
  • RQ3本フレームワークは、オンラインレビューに見られる口語的・口語的ペルシャ語語彙をどの程度効果的に処理できるか?
  • RQ4特徴選択は、ペルシャ語のセンチメント分析において次元削減と分類速度の向上にどのような影響を与えるか?
  • RQ5提案されたモデルは、スマートフォンレビュー以外のペルシャ語テキスト分野にも一般化可能か?

主な発見

  • 提案された特徴選択手法は、ベースライン手法と比較して、ペルシャ語レビューのセンチメント分類精度を顕著に向上させた。
  • 語彙還元は語形のばらつきを効果的に低減し、特徴表現と分類性能の向上に寄与した。
  • 本フレームワークは、低リソースNLPにおける主要な課題である口語的・口語的ペルシャ語の処理において、頑健性を示した。
  • 特徴選択による次元削減は、処理速度の向上とモデル効率の改善をもたらした。
  • 手動で整備したペルシャ語スマートフォンレビューのデータセット上で、モデルは優れた性能を発揮し、実用的応用の有効性が裏付けられた。
  • 結果から、本手法のフレームワークが、ペルシャ語の低リソースセンチメント分析環境において有効であることが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。