[Paper Review] Challenges in Translation of Emotions in Multilingual User-Generated Content: Twitter as a Case Study
This study investigates the reliability of neural machine translation (NMT) systems in preserving emotional content in multilingual user-generated content, focusing on Twitter. It identifies linguistic features like contronyms, diacritics, idioms, code-switching, and negation as key causes of emotion mistranslation, showing that standard automatic metrics like BLEU and METEOR fail to detect sentiment flips, even when emotion polarity is completely reversed.
Although emotions are universal concepts, transferring the different shades of emotion from one language to another may not always be straightforward for human translators, let alone for machine translation systems. Moreover, the cognitive states are established by verbal explanations of experience which is shaped by both the verbal and cultural contexts. There are a number of verbal contexts where expression of emotions constitutes the pivotal component of the message. This is particularly true for User-Generated Content (UGC) which can be in the form of a review of a product or a service, a tweet, or a social media post. Recently, it has become common practice for multilingual websites such as Twitter to provide an automatic translation of UGC to reach out to their linguistically diverse users. In such scenarios, the process of translating the user's emotion is entirely automatic with no human intervention, neither for post-editing nor for accuracy checking. In this research, we assess whether automatic translation tools can be a successful real-life utility in transferring emotion in user-generated multilingual data such as tweets. We show that there are linguistic phenomena specific of Twitter data that pose a challenge in translation of emotions in different languages. We summarise these challenges in a list of linguistic features and show how frequent these features are in different language pairs. We also assess the capacity of commonly used methods for evaluating the performance of an MT system with respect to the preservation of emotion in the source text.
Motivation & Objective
- To assess whether automatic machine translation (MT) systems can accurately transfer emotional content in multilingual user-generated content (UGC), particularly on Twitter.
- To identify specific linguistic features in tweets that commonly lead to mistranslation of emotions across different language pairs.
- To evaluate the effectiveness of standard automatic MT evaluation metrics—such as BLEU and METEOR—in detecting emotion-preserving errors in UGC.
- To investigate whether fine-grained emotions (e.g., joy, anger, fear) are preserved during translation, beyond just sentiment polarity.
- To highlight the limitations of current MT evaluation practices in capturing affective message distortion in informal, context-rich text like tweets.
Proposed method
- Collected multilingual Twitter datasets previously annotated for four emotions: joy, fear, aggression, and anger, sourced from shared tasks in emotion and aggression detection.
- Used Google Translate API to automatically translate tweets from Arabic, Spanish, and other languages into English, simulating real-world multilingual platform translation workflows.
- Conducted qualitative and quantitative analysis of six linguistic features—contronyms, diacritics, idiomatic expressions, dialectal code-switching, negation, and punctuation—common in tweets that disrupt emotion transfer.
- Compared machine-translated outputs against human reference translations to compute standard MT evaluation metrics (BLEU and METEOR) on a 100-tweet dataset of emotion-mistranslated examples.
- Analyzed the correlation between automatic metric scores and human judgment of sentiment preservation, focusing on cases where emotion polarity was flipped.
- Proposed that future MT evaluation should incorporate sentiment-aware metrics to better reflect affective message fidelity in UGC.
Experimental results
Research questions
- RQ1Are there specific linguistic features in tweets that lead to mistranslation of emotions in multilingual NMT systems?
- RQ2Do these linguistic features affect different language pairs equally in terms of emotion distortion?
- RQ3Can traditional automatic MT evaluation metrics (e.g., BLEU, METEOR) adequately detect mistranslations of sentiment in UGC, particularly when emotion polarity is flipped?
- RQ4To what extent do standard metrics overestimate translation quality when the emotional content of a tweet is completely inverted?
Key findings
- Linguistic features such as negation, contronyms, and idiomatic expressions significantly disrupt emotion transfer, with negation causing complete sentiment reversal in 15% of analyzed cases.
- The BLEU score for mistranslated emotion examples averaged 0.60, and METEOR scored 0.45, indicating high metric scores despite complete emotion polarity flips, showing poor correlation with human judgment.
- A single mistranslation—such as replacing 'happiness' with 'anger'—received a BLEU score of 0.76, demonstrating that lexical similarity metrics fail to penalize semantic distortion of affective meaning.
- The study found that METEOR, despite incorporating semantic features, still failed to detect sentiment flips, such as in a case where negation was omitted, flipping anger to joy, which received a METEOR score of 0.61.
- The results indicate that standard MT evaluation metrics are inadequate for assessing sentiment preservation in UGC, especially in emotionally charged, informal text like tweets.
- The research concludes that current automatic evaluation methods overestimate translation quality when emotion is distorted, calling for sentiment-aware metrics in MT evaluation for affective text.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.