[논문 리뷰] Tweets Sentiment Analysis via Word Embeddings and Machine Learning Techniques
이 논문은 2019년 총선 관련 실시간 트위터 데이터를 대상으로 단어 임베딩 기반의 문맥적 특징 표현을 위한 word2vec과 분류를 위한 랜덤 포레스트를 활용한 감성 분석 프레임워크를 제안한다. 기존의 BOW 및 TF-IDF와 같은 전통적 방법에 비해 분산 단어 임베딩을 통해 텍스트 내 의미 관계를 포착함으로써 정확도를 크게 향상시킨다.
Sentiment analysis of social media data consists of attitudes, assessments, and emotions which can be considered a way human think. Understanding and classifying the large collection of documents into positive and negative aspects are a very difficult task. Social networks such as Twitter, Facebook, and Instagram provide a platform in order to gather information about peoples sentiments and opinions. Considering the fact that people spend hours daily on social media and share their opinion on various different topics helps us analyze sentiments better. More and more companies are using social media tools to provide various services and interact with customers. Sentiment Analysis (SA) classifies the polarity of given tweets to positive and negative tweets in order to understand the sentiments of the public. This paper aims to perform sentiment analysis of real-time 2019 election twitter data using the feature selection model word2vec and the machine learning algorithm random forest for sentiment classification. Word2vec with Random Forest improves the accuracy of sentiment analysis significantly compared to traditional methods such as BOW and TF-IDF. Word2vec improves the quality of features by considering contextual semantics of words in a text hence improving the accuracy of machine learning and sentiment analysis.
연구 동기 및 목표
- 대규모 실시간 소셜미디어 데이터, 특히 2019년 총선 기간 동안의 공공의 감성을 분류하는 데 도전하는 것.
- 기존의 BOW 및 TF-IDF와 같은 전통적 특징 추출 방법을 word2vec 임베딩으로 대체하여 감성 분류 정확도를 향상시키는 것.
- word2vec과 랜덤 포레스트 알고리즘을 조합하여 트위터에서의 감성 극성 탐지 성능을 평가하는 것.
- word2vec에서 유도된 문맥적 의미 특징이 감성 분석 작업에서 기계 학습 성능을 향상시킨다는 것을 입증하는 것.
제안 방법
- word2vec을 사용하여 훈련 코퍼스에서 유의미한 문맥적 의미를 반영하는 조밀하고 분산된 단어 벡터 표현을 생성한다.
- word2vec 모델은 2019년 총선 관련 트위터 데이터로 구성된 대규모 코퍼스에서 훈련되어 의미 있는 단어 임베딩을 학습한다.
- word2vec에서 유도된 특징 벡터를 랜덤 포레스트 분류기의 입력으로 사용하여 감성 극성 예측을 수행한다.
- 랜덤 포레스트 알고리즘을 적용하여 학습된 단어 임베딩 기반으로 트위터를 긍정 또는 부정 감성으로 분류한다.
- 특징 선택 모델을 word2vec과 통합하여 입력 특징의 품질을 향상시키고 차원을 감소시킨다.
- word2vec + 랜덤 포레스트 파이프라인의 성능을 BOW 및 TF-IDF와 같은 기준선 방법과 비교한다.
실험 결과
연구 질문
- RQ1word2vec 기반 특징 표현은 BOW 및 TF-IDF와 같은 전통적 방법에 비해 감성 분류 정확도를 얼마나 향상시키는가?
- RQ2word2vec과 랜덤 포레스트를 조합함으로써 트위터 데이터에서 감성 분석 성능이 얼마나 향상되는가?
- RQ3문맥적 단어 임베딩은 짧고 비공식적인 소셜미디어 텍스트에서 감성과 관련된 의미를 효과적으로 포착할 수 있는가?
- RQ4특징 선택은 word2vec 임베딩의 품질과 이후 분류 정확도에 어떤 영향을 미치는가?
주요 결과
- 제안된 word2vec과 랜덤 포레스트를 활용한 방법은 BOW 및 TF-IDF와 같은 기존 방법에 비해 감성 분류 정확도에서 뚜렷한 향상을 이룬다.
- word2vec은 문맥적 의미를 코딩함으로써 기계 학습 모델의 입력에 대한 구분 능력을 향상시켜 특징 품질을 향상시킨다.
- word2vec과 랜덤 포레스트의 통합은 트위터 감성 분류에서 더 나은 일반화 및 강건성을 이끌어낸다.
- 문맥적 단어 임베딩의 사용은 BOW 가정에 대한 의존도를 감소시켜 보다 세밀한 감성 탐지 결과를 도출한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.