Skip to main content
QUICK REVIEW

[论文解读] Sentiment Analysis of Persian Language: Review of Algorithms, Approaches and Datasets

Ali Nazarizadeh, Touraj Banirostam|arXiv (Cornell University)|Nov 9, 2022
Sentiment Analysis and Opinion Mining被引用 7
一句话总结

本综述论文整合了2018至2022年间40项针对波斯语的最新情感分析方法,评估了机器学习与深度学习模型(包括BERT、LSTM和Bi-LSTM)在12个基准数据集上的表现。研究发现,基于Transformer的模型如BERT以及RNN变体如Bi-LSTM表现最为准确,为波斯语自然语言处理研究提供了方法、数据集与性能指标的全面分析。

ABSTRACT

Sentiment analysis aims to extract people's emotions and opinion from their comments on the web. It widely used in businesses to detect sentiment in social data, gauge brand reputation, and understand customers. Most of articles in this area have concentrated on the English language whereas there are limited resources for Persian language. In this review paper, recent published articles between 2018 and 2022 in sentiment analysis in Persian Language have been collected and their methods, approach and dataset will be explained and analyzed. Almost all the methods used to solve sentiment analysis are machine learning and deep learning. The purpose of this paper is to examine 40 different approach sentiment analysis in the Persian Language, analysis datasets along with the accuracy of the algorithms applied to them and also review strengths and weaknesses of each. Among all the methods, transformers such as BERT and RNN Neural Networks such as LSTM and Bi-LSTM have achieved higher accuracy in the sentiment analysis. In addition to the methods and approaches, the datasets reviewed are listed between 2018 and 2022 and information about each dataset and its details are provided.

研究动机与目标

  • 系统回顾并分析2018至2022年间波斯语情感分析的最新进展。
  • 评估各类机器学习与深度学习算法在波斯语文本上的性能表现。
  • 整理并比较公开可用的波斯语情感分析数据集。
  • 识别各项研究在方法论与数据整理方面的优势、劣势及发展趋势。
  • 为低资源语言情感分析研究(特别是波斯语)提供参考框架。

提出的方法

  • 本研究对2018至2022年间发表的40篇经同行评审的波斯语情感分析论文进行了系统性回顾。
  • 将方法分类为机器学习与深度学习范式,重点聚焦于Transformer模型(如BERT)与循环网络(如LSTM、Bi-LSTM)。
  • 基于规模、来源、标注质量与领域特异性对数据集进行分析。
  • 提取并比较各类模型在准确率、精确率、召回率与F1分数等性能指标上的表现。
  • 评估各研究中使用的模型架构、预处理技术与数据增强策略。
  • 开展对比分析,以识别最有效的算法与数据集特征。

实验结果

研究问题

  • RQ1在2018至2022年期间,哪些机器学习与深度学习模型在波斯语情感分析任务中实现了最高准确率?
  • RQ2基于Transformer的模型(如BERT)与基于RNN的模型(如LSTM与Bi-LSTM)在波斯语文本上的性能表现如何比较?
  • RQ32018至2022年间公开可用的波斯语情感分析数据集的关键特征与局限性是什么?
  • RQ4波斯语情感分析研究中最常见的预处理与特征工程技术有哪些?
  • RQ5当前波斯语自然语言处理研究领域(特别是情感分析方向)存在哪些趋势与研究空白?

主要发现

  • 基于Transformer的模型(如BERT)在波斯语文本的情感分类任务中实现了最高准确率。
  • 循环神经网络(尤其是Bi-LSTM)表现出色,属于表现最佳的模型之一。
  • 本研究识别出2018至2022年间发布的12个不同的波斯语情感分析数据集,其规模与领域覆盖范围各不相同。
  • 多数研究依赖于人工标注的监督学习数据集,凸显了对人工标注数据的依赖。
  • 尽管已取得进展,但目前仍缺乏大规模、多样化且公开可用的波斯语情感分析数据集。
  • 综述揭示出一个持续趋势:深度学习模型(尤其是微调的预训练Transformer)被广泛采用,以提升性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。