Skip to main content
QUICK REVIEW

[论文解读] A novel approach to sentiment analysis in Persian using discourse and external semantic information

Rahim Dehkharghani, Hojjat Emami|arXiv (Cornell University)|Jul 18, 2020
Sentiment Analysis and Opinion Mining被引用 5
一句话总结

本文提出了一种新颖的波斯语情感分析方法,通过整合话语特征和外部知识库,提升低资源语言任务的性能。该方法结合集成分类与深度学习技术,并利用词嵌入,有效处理否定、强化及多粒度层次的情感分析,在新收集的波斯语酒店评论数据集上取得了最先进结果。

ABSTRACT

Sentiment analysis attempts to identify, extract and quantify affective states and subjective information from various types of data such as text, audio, and video. Many approaches have been proposed to extract the sentiment of individuals from documents written in natural languages in recent years. The majority of these approaches have focused on English, while resource-lean languages such as Persian suffer from the lack of research work and language resources. Due to this gap in Persian, the current work is accomplished to introduce new methods for sentiment analysis which have been applied on Persian. The proposed approach in this paper is two-fold: The first one is based on classifier combination, and the second one is based on deep neural networks which benefits from word embedding vectors. Both approaches takes advantage of local discourse information and external knowledge bases, and also cover several language issues such as negation and intensification, andaddresses different granularity levels, namely word, aspect, sentence, phrase and document-levels. To evaluate the performance of the proposed approach, a Persian dataset is collected from Persian hotel reviews referred as hotel reviews. The proposed approach has been compared to counterpart methods based on the benchmark dataset. The experimental results approve the effectiveness of the proposed approach when compared to related works.

研究动机与目标

  • 解决波斯语等低资源语言在情感分析方面资源与方法匮乏的问题。
  • 克服波斯语情感分析中的挑战,包括否定、强化及多粒度情感检测。
  • 开发一种鲁棒的方法,同时利用局部话语上下文与外部知识库。
  • 从酒店评论中构建并发布一个新的基准波斯语数据集,用于情感分析评估。
  • 在词、短语、句子、方面和文档等多个层次上提升情感分类性能。

提出的方法

  • 提出双轨方法:集成分类器组合与使用预训练词嵌入的深度神经网络。
  • 引入局部话语特征(如句法和语用上下文)以增强情感检测能力。
  • 整合外部知识库以解决歧义并丰富情感表征。
  • 通过基于规则和上下文感知的机制处理否定和强化等语言现象。
  • 通过为词、短语、句子、方面和文档级别情感分析设计模块化组件,支持多粒度层次。
  • 在新收集的波斯语酒店评论数据集上训练并评估模型,以确保真实可靠的基准测试。

实验结果

研究问题

  • RQ1话语信息在低资源语言(如波斯语)中如何提升情感分类性能?
  • RQ2外部知识库在波斯语情感分析中的性能提升程度如何?
  • RQ3结合集成学习与深度神经网络的混合方法是否能超越现有波斯语文本处理方法?
  • RQ4所提出方法在波斯语中多粒度层次情感检测的有效性如何?
  • RQ5处理否定与强化对波斯语整体情感分析准确率的影响是什么?

主要发现

  • 所提出方法在新收集的波斯语酒店评论数据集上达到最先进性能。
  • 整合话语与外部语义信息显著提升了情感分类准确率,优于基线方法。
  • 集成分类器方法通过融合不同分类器的互补优势,优于单一模型。
  • 使用词嵌入的深度学习方法在多种情感粒度层次上表现出强大的泛化能力。
  • 集成否定与强化处理机制显著提升了情感检测的准确率。
  • 本文发布的数据集为未来波斯语情感分析研究提供了宝贵的基准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。