Skip to main content
QUICK REVIEW

[論文レビュー] Challenges in Translation of Emotions in Multilingual User-Generated Content: Twitter as a Case Study

Hadeel Saadany, Constantin Orǎsan|arXiv (Cornell University)|Jun 20, 2021
Natural Language Processing Techniques被引用数 8
ひとこと要約

本研究は、多言語のユーザーゲネレーテッドコンテンツ(UGC)、特にTwitterにおける感情的コンテンツの保存に関するニューラル機械翻訳(NMT)システムの信頼性を調査する。対照語、変換記号、慣用句、コードスイッチング、否定文などの言語的特徴が感情の誤訳の主な原因であると特定し、BLEU や METEOR といった標準的な自動評価指標が、感情極性が完全に反転している場合でも感情の反転を検出できないことを示している。

ABSTRACT

Although emotions are universal concepts, transferring the different shades of emotion from one language to another may not always be straightforward for human translators, let alone for machine translation systems. Moreover, the cognitive states are established by verbal explanations of experience which is shaped by both the verbal and cultural contexts. There are a number of verbal contexts where expression of emotions constitutes the pivotal component of the message. This is particularly true for User-Generated Content (UGC) which can be in the form of a review of a product or a service, a tweet, or a social media post. Recently, it has become common practice for multilingual websites such as Twitter to provide an automatic translation of UGC to reach out to their linguistically diverse users. In such scenarios, the process of translating the user's emotion is entirely automatic with no human intervention, neither for post-editing nor for accuracy checking. In this research, we assess whether automatic translation tools can be a successful real-life utility in transferring emotion in user-generated multilingual data such as tweets. We show that there are linguistic phenomena specific of Twitter data that pose a challenge in translation of emotions in different languages. We summarise these challenges in a list of linguistic features and show how frequent these features are in different language pairs. We also assess the capacity of commonly used methods for evaluating the performance of an MT system with respect to the preservation of emotion in the source text.

研究の動機と目的

  • 自動機械翻訳(MT)システムが、特にTwitter上で発生する多言語のユーザーゲネレーテッドコンテンツ(UGC)における感情的コンテンツを正確に伝えることができるかどうかを評価すること。
  • 異なる言語ペア間で感情の誤訳を引き起こす要因となる、ツイートに一般的に見られる特定の言語的特徴を同定すること。
  • BLEU や METEOR などの標準的な自動MT評価指標が、UGCにおける感情保持エラーを検出する効果性を評価すること。
  • 感情極性を超えて、喜び、怒り、恐怖といった微細な感情(例:喜び、攻撃的、恐怖)が翻訳中に保持されるかどうかを調査すること。
  • 現在のMT評価手法が、感情的メッセージの歪みを捉えることができない点、特に文脈が豊富で非公式なテキスト(例:ツイート)において顕著であることを強調すること。

提案手法

  • 感情検出共有タスクから得られた、喜び、恐怖、攻撃的、怒りの4つの感情を事前にアノテート済みの多言語のツイートデータセットを収集。
  • Google Translate API を用いて、アラビア語、スペイン語、その他の言語から英語への自動翻訳を実施し、実世界の多言語プラットフォームの翻訳ワークフローを模擬。
  • ツイートに一般的に見られる6つの言語的特徴(対照語、変換記号、慣用句、方言的コードスイッチング、否定、句読点)が感情伝達に与える影響を定性的かつ定量的に分析。
  • 100件の感情誤訳例から成るデータセットを対象に、機械翻訳出力と人間による基準翻訳を比較し、標準的なMT評価指標(BLEU と METEOR)を計算。
  • 自動指標スコアと人間による感情保持の判断との相関関係を分析し、特に感情極性が反転したケースに注目。
  • 今後のMT評価には、感情を意識した指標を組み込むべきであると提言。

実験結果

リサーチクエスチョン

  • RQ1ツイートに見られる特定の言語的特徴が、多言語NMTシステムにおける感情誤訳を引き起こすか?
  • RQ2これらの言語的特徴は、異なる言語ペアにおいて感情歪みの度合いに同様の影響を及ぼすか?
  • RQ3従来の自動MT評価指標(例:BLEU、METEOR)は、UGCにおける感情誤訳、特に感情極性が反転している場合に十分に検出できるか?
  • RQ4感情の内容が完全に逆転している場合でも、標準的な指標が翻訳品質を過大評価する程度はどの程度か?

主な発見

  • 否定、対照語、慣用句といった言語的特徴が感情伝達を著しく妨げており、否定は分析対象の15%のケースで完全な感情極性の反転を引き起こしている。
  • 誤訳例におけるBLEUスコアの平均は0.60、METEORスコアは0.45であり、感情極性が完全に反転しているにもかかわらず高いスコアを示しており、人間の判断とは相関が低いことが判明。
  • たった1つの誤訳例(例:「喜び」が「怒り」に置き換えられたもの)でも、BLEUスコアが0.76に達しており、語彙的類似度指標が感情的意味の意味的歪みを罰しないことが示された。
  • METEORでさえも、否定が省かれて怒りが喜びに反転したケース(METEORスコア0.61)を検出できず、感情の反転を検出できないことが判明。
  • 標準的なMT評価指標は、特に感情的で非公式なテキスト(例:ツイート)における感情的メッセージの歪みを評価するのに不適切であることが示された。
  • 本研究の結論として、現在の自動評価手法は感情が歪められている場合に翻訳品質を過大評価しており、感情的テキストのMT評価には感情を意識した指標の導入が不可欠であると結論づけた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。