Skip to main content
QUICK REVIEW

[논문 리뷰] Phonetic-enriched Text Representation for Chinese Sentiment Analysis with Reinforcement Learning

Haiyun Peng, Yukun Ma|arXiv (Cornell University)|2019. 01. 23.
Sentiment Analysis and Opinion Mining참고 문헌 55인용 수 5
한 줄 요약

이 논문은 강화학습 기반의 DISA 네트워크를 통해 깊이 있는 음소 철자법과 억양 변동을 활용하여 중국어 감성 분석을 위한 새로운 청각적 특징을 통합한 텍스트 표현 방식을 제안한다. 학습된 음성 특징과 추출된 음성 특징을 텍스트 및 시각 모odalities와 융합하여, 음성 특징의 해석을 통해 감성 분류 정확도를 크게 향상시켜 다섯 개인 중국어 감성 분석 데이터셋에서 최신 기술 수준의 성능을 달성한다.

ABSTRACT

The Chinese pronunciation system offers two characteristics that distinguish it from other languages: deep phonemic orthography and intonation variations. We are the first to argue that these two important properties can play a major role in Chinese sentiment analysis. Particularly, we propose two effective features to encode phonetic information. Next, we develop a Disambiguate Intonation for Sentiment Analysis (DISA) network using a reinforcement network. It functions as disambiguating intonations for each Chinese character (pinyin). Thus, a precise phonetic representation of Chinese is learned. Furthermore, we also fuse phonetic features with textual and visual features in order to mimic the way humans read and understand Chinese text. Experimental results on five different Chinese sentiment analysis datasets show that the inclusion of phonetic features significantly and consistently improves the performance of textual and visual representations and outshines the state-of-the-art Chinese character level representations.

연구 동기 및 목표

  • 중국어 음성학, 특히 깊이 있는 음소 철자법과 억양 변동이 감성 분석에서 간과된 역할을 다루는 것.
  • 각 파인진에 대해 억양을 학습하고 해석하는 방법을 개발하여 감성 분류 성능을 향상시키는 것.
  • 음성 특징을 텍스트 및 시각적 표현과 융합하여 인간의 다중 모odal 이해 방식을 모방하는 것.
  • 음성 정보가 중국어 감성 분석에서 텍스트 및 시각적 표현을 향상시킬 수 있음을 입증하는 것.

제안 방법

  • 실제 클립에서 유래한 오디오 기반 특징과 파인진 어휘집에서 학습된 파인진 토큰 임베딩을 각각 억양 유무와 함께 두 가지 유형의 음성 특징으로 제안한다.
  • 액터-크리틱 아키텍처를 갖춘 강화학습 프레임워크를 사용하여 DISA(감성 분석을 위한 억양 해석) 네트워크를 설계한다.
  • 액터 네트워크는 각 파인진에 대해 다섯 가지 억양 중 하나를 선택하고, 크리틱 네트워크는 파인진 시퀀스를 인코딩하고 감성 예측을 계산한다.
  • 정책 네트워크는 감성 분류 성능에 기반한 지연 보상에 따라 업데이트되고, 크리틱 네트워크는 교차 엔트로피 손실을 통해 훈련된다.
  • 음성 특징을 텍스트 및 시각적 특징과 융합하여 감성 분류를 위한 다중 모달 표현을 생성한다.
  • t-SNE 시각화를 통해 학습된 임베딩이 음성 유사성과 의미 관계를 모두 포착하고 있음을 검증한다.
Figure 1 : Original input bitmaps (upper row) and reconstructed output bitmaps (lower row).
Figure 1 : Original input bitmaps (upper row) and reconstructed output bitmaps (lower row).

실험 결과

연구 질문

  • RQ1특히 억양 변동을 포함한 음성 정보가 텍스트 및 시각적 특징만으로는 중국어 감성 분석을 향상시키는 데 기여할 수 있는가?
  • RQ2강화학습 모델이 맥락 속에서 각 파인진에 대해 올바른 억양을 얼마나 효과적으로 해석할 수 있는가?
  • RQ3음성 특징이 다중 모달 중국어 감성 분석에서 텍스트 및 시각적 표현을 어느 정도 향상시키는가?
  • RQ4학습된 음성 임베딩이 중국어 파인진에서 음성 유사성과 의미를 모두 포착할 수 있는가?

주요 결과

  • 추출된(Ex04) 및 학습된(PW) 음성 특징의 조합(Ex04+PW)이 Ex04만을 사용할 때보다 다섯 개의 데이터셋 평균 1.43% 향상된 성능을 달성했다.
  • 텍스트 및 시각 모달리티와 음성 특징을 융합한(T+P) 모델이 기존 최신 기술 수준의 방법보다 통계적으로 유의미한 성능 향상을 보였다.
  • t-SNE 시각화 결과, Ex04+PW 임베딩이 음성 유사성(예: 유사한 모음)과 의미 유사성(예: 'Niu2'와 'Nai3'가 '우유'를 의미함)을 모두 포착하고 있음을 확인했다.
  • 융합된 T+P 표현은 의미 관계(예: 'Mu4'와 'Yu4'가 '목욕'을 의미함)와 음성 관계(예: 'Huan2'와 'Huan2'가 동음이의어임)를 모두 효과적으로 인코딩했다.
  • DISA 네트워크는 억양을 효과적으로 해석하여 감성 극성의 모호성을 해결했다(예: 'hǎochi'(맛있다) 대비 'hàochi'(탐욕스럽다)).
  • 이 연구는 깊이 있는 음소 철자법과 억양이 텍스트만으로는 포착되지 않는 의미적 및 감성적 신호를 지닌다는 것을 입증했다.
Figure 2: DISA model structure for tone selection. $C_{m}$ stands for the $m$ th Chinese character in a sentence. $P_{m}$ denotes the pinyin for $m$ th character without the tones. $P_{m}n$ represents the pinyin for $m$ th character with its $n$ th tone. $F_{m}n$ is the feature/embedding vector for
Figure 2: DISA model structure for tone selection. $C_{m}$ stands for the $m$ th Chinese character in a sentence. $P_{m}$ denotes the pinyin for $m$ th character without the tones. $P_{m}n$ represents the pinyin for $m$ th character with its $n$ th tone. $F_{m}n$ is the feature/embedding vector for

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.