Skip to main content
QUICK REVIEW

[논문 리뷰] Sentiment analysis in tweets: an assessment study from classical to modern text representation models

Sérgio Barreto, Ricardo Moura|arXiv (Cornell University)|2021. 05. 29.
Sentiment Analysis and Opinion Mining참고 문헌 51인용 수 5
한 줄 요약

이 연구는 트윗의 감성 분석을 위한 고전적 및 현대적 텍스트 표현 모델을 평가하며, 정적 임베딩(예: TF-IDF, Word2Vec)과 문맥 기반 모델(예: BERT, RoBERTa, BERTweet)을 비교한다. BERTweet는 트윗에 최적화된 BERT의 변종으로 감성 데이터셋에 피니튜닝된 모델로, 특히 로지스틱 회귀나 MLP 분류기와 조합했을 때 22개의 다양한 데이터셋에서 원본 및 피니튜닝된 모델을 모두 능가하는 최고의 성능을 기록한다.

ABSTRACT

With the growth of social medias, such as Twitter, plenty of user-generated data emerge daily. The short texts published on Twitter -- the tweets -- have earned significant attention as a rich source of information to guide many decision-making processes. However, their inherent characteristics, such as the informal, and noisy linguistic style, remain challenging to many natural language processing (NLP) tasks, including sentiment analysis. Sentiment classification is tackled mainly by machine learning-based classifiers. The literature has adopted word representations from distinct natures to transform tweets to vector-based inputs to feed sentiment classifiers. The representations come from simple count-based methods, such as bag-of-words, to more sophisticated ones, such as BERTweet, built upon the trendy BERT architecture. Nevertheless, most studies mainly focus on evaluating those models using only a small number of datasets. Despite the progress made in recent years in language modelling, there is still a gap regarding a robust evaluation of induced embeddings applied to sentiment analysis on tweets. Furthermore, while fine-tuning the model from downstream tasks is prominent nowadays, less attention has been given to adjustments based on the specific linguistic style of the data. In this context, this study fulfils an assessment of existing language models in distinguishing the sentiment expressed in tweets by using a rich collection of 22 datasets from distinct domains and five classification algorithms. The evaluation includes static and contextualized representations. Contexts are assembled from Transformer-based autoencoder models that are also fine-tuned based on the masked language model task, using a plethora of strategies.

연구 동기 및 목표

  • 트윗의 감성 분석에서 고전적 및 현대적 텍스트 표현 모델의 성능을 평가하기 위해.
  • 트윗의 비공식적이고 노이즈가 많은 언어 스타일을 다룰 때 정적 임베딩과 문맥 기반 임베딩 간의 비교를 조사하기 위해.
  • Transformer 기반 모델을 도메인 특화 감성 데이터셋에 비해 일반적인 트윗 코퍼스에 대해 피니튜닝했을 때의 영향을 평가하기 위해.
  • 트윗 감성 분류에 최적의 언어 모델과 분류기 조합을 규명하기 위해.
  • 일반 트윗에 대해 사전 훈련하는 것과 감성 레이블이 부여된 데이터에 대해 피니튜닝하는 것이 트윗 감성 분석에서 더 나은 결과를 낳는지 확인하기 위해.

제안 방법

  • 다양한 도메인을 포함하는 22개의 다양한 영어 트윗 데이터셋을 평가하여 탄탄한 결과를 확보했다.
  • 정적 표현(백오프워즈, TF-IDF, Word2Vec, Fast-Text)과 문맥 기반 모델(BERT, RoBERTa, BERTweet)을 비교했다.
  • 다양한 크기의 비라벨 영어 트윗(0.5K에서 1.5M까지)을 사용해 마스킹 언어 모델링을 통해 Transformer 기반 모델을 피니튜닝했다.
  • 다양한 학습 패러다임을 평가하기 위해 다섯 가지 분류기(로지스틱 회귀, SVM, MLP, 랜덤 포레스트, 나이브 베이즈)를 적용했다.
  • 목표 데이터셋과 외부 감성 데이터셋을 모두 사용한 하이브리드 피니튜닝을 포함해, 일반 트윗과 감성 레이블이 부여된 트윗에 대해 피니튜닝하는 것을 비교한 추론 연구를 수행했다.
  • 대규모 트윗 코퍼스에 사전 훈련된 BERTweet(트윗 전용 BERT 변종)를 사용해 도메인 특화 사전 훈련과 피니튜닝의 이점을 평가했다.

실험 결과

연구 질문

  • RQ1TF-IDF, Word2Vec 등의 정적 텍스트 표현 방식과 BERT, RoBERTa, BERTweet 등의 문맥 기반 표현 방식은 트윗 감성 분류에서 어떻게 비교되는가?
  • RQ2대규모 비라벨 영어 트윗에 대해 사전 훈련된 Transformer 기반 모델을 피니튜닝하면 감성 분류 성능이 향상되는가?
  • RQ3감성 레이블이 부여된 데이터셋에 대해 피니튜닝하는 것이 일반 트윗 코퍼스에 대해 피니튜닝하는 것보다 감성 분석에서 더 효과적인가?
  • RQ4언어 모델과 분류기의 조합 중 어느 것이 트윗 감성 분류에서 가장 높은 성능을 낳는가?
  • RQ5피니튜닝 데이터셋의 크기가 BERT, RoBERTa, BERTweet와 같은 사전 훈련된 모델의 성능에 어떻게 영향을 미치는가?

주요 결과

  • BERTweet는 목표 데이터셋과 대규모 외부 감성 데이터셋을 조합해 피니튜닝한 BERTweet_22Dt 설정에서 22개 모든 데이터셋에서 최고의 전체 성능을 기록했다.
  • BERTweet는 단지 5,000개의 피니튜닝 트윗만으로도 다른 모든 모델을 능가했으며, 이는 트윗 데이터에 사전 훈련된 모델가 더 작은 피니튜닝 데이터로도 높은 성능을 달성할 수 있음을 시사한다.
  • 감성 레이블이 부여된 데이터셋에 대해 피니튜닝하는 것이 일반 트윗에 대해 피니튜닝하는 것보다 성능 향상이 일관되게 이루어졌으며, BERTweet_22Dt는 모든 설정에서 최고의 F1 스코어를 기록했다.
  • BERTweet와 로지스틱 회귀(LR) 및 다층 퍼셉트론(MLP) 분류기의 조합이 가장 뛰어난 전체 결과를 낳았으며, 다른 모델-분류기 조합을 모두 능가했다.
  • RoBERTa와 BERT는 각각 250,000개와 50,000개의 트윗으로 피니튜닝했을 때 성능이 향상되었지만, 더 큰 피니튜닝 세트를 사용하더라도 여전히 BERTweet에 뒤지지 않았다.
  • 모델에 따라 피니튜닝 전략이 달랐다: BERT는 목표 데이터셋만으로도 최고의 성능를 보였지만, RoBERTa와 BERTweet는 목표 데이터셋과 외부 감성 데이터셋을 함께 사용할 때 성능 향상이 뚜렷했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.