[论文解读] Challenges in Translation of Emotions in Multilingual User-Generated Content: Twitter as a Case Study
本研究调查了神经机器翻译(NMT)系统在多语言用户生成内容中保持情感内容的可靠性,重点关注推特平台。研究识别出诸如反义词、变音符号、习语、语码转换和否定等语言特征是情感误译的主要原因,表明标准自动评估指标如BLEU和METEOR无法检测到情感极性反转,即使情感极性被完全颠倒。
Although emotions are universal concepts, transferring the different shades of emotion from one language to another may not always be straightforward for human translators, let alone for machine translation systems. Moreover, the cognitive states are established by verbal explanations of experience which is shaped by both the verbal and cultural contexts. There are a number of verbal contexts where expression of emotions constitutes the pivotal component of the message. This is particularly true for User-Generated Content (UGC) which can be in the form of a review of a product or a service, a tweet, or a social media post. Recently, it has become common practice for multilingual websites such as Twitter to provide an automatic translation of UGC to reach out to their linguistically diverse users. In such scenarios, the process of translating the user's emotion is entirely automatic with no human intervention, neither for post-editing nor for accuracy checking. In this research, we assess whether automatic translation tools can be a successful real-life utility in transferring emotion in user-generated multilingual data such as tweets. We show that there are linguistic phenomena specific of Twitter data that pose a challenge in translation of emotions in different languages. We summarise these challenges in a list of linguistic features and show how frequent these features are in different language pairs. We also assess the capacity of commonly used methods for evaluating the performance of an MT system with respect to the preservation of emotion in the source text.
研究动机与目标
- 评估自动机器翻译(MT)系统是否能准确传递多语言用户生成内容(UGC)中的情感内容,特别是在推特平台上的表现。
- 识别推文中常见导致情感在不同语言对之间误译的具体语言特征。
- 评估标准自动MT评估指标(如BLEU和METEOR)在检测UGC中情感保持错误方面的有效性。
- 调查细微情感(如喜悦、愤怒、恐惧)在翻译过程中是否得以保留,而不仅限于情感极性。
- 指出当前MT评估实践在捕捉非正式、语境丰富的文本(如推文)中情感信息失真方面的局限性。
提出的方法
- 收集了先前在情绪和攻击性检测共享任务中已标注四种情绪(喜悦、恐惧、攻击性、愤怒)的多语言推特数据集。
- 使用Google Translate API将阿拉伯语、西班牙语及其他语言的推文自动翻译成英文,模拟真实世界多语言平台的翻译工作流程。
- 对推文中常见的六种语言特征(反义词、变音符号、习语表达、方言语码转换、否定和标点符号)进行定性和定量分析,这些特征常导致情感传递受阻。
- 将机器翻译结果与人工参考翻译进行对比,针对100个情感误译示例的数据集计算标准MT评估指标(BLEU和METEOR)。
- 分析自动指标得分与人类对情感保留判断之间的相关性,重点关注情感极性被反转的案例。
- 建议未来MT评估应引入情感感知指标,以更好地反映UGC中情感信息的保真度。
实验结果
研究问题
- RQ1推文中是否存在特定语言特征,导致多语言NMT系统中情感误译?
- RQ2这些语言特征在不同语言对中是否以相同程度影响情感失真?
- RQ3传统自动MT评估指标(如BLEU、METEOR)能否充分检测UGC中的情感误译,特别是当情感极性被反转时?
- RQ4当推文的情感内容被完全颠倒时,标准指标在多大程度上高估了翻译质量?
主要发现
- 否定、反义词和习语表达等语言特征显著干扰情感传递,其中否定在15%的分析案例中导致情感极性完全反转。
- 情感误译示例的BLEU平均得分为0.60,METEOR得分为0.45,表明尽管情感极性完全颠倒,指标得分仍较高,与人类判断的相关性较差。
- 一次单一误译——如将“happiness”替换为“anger”——获得了0.76的BLEU得分,表明词汇相似性指标无法惩罚情感意义的语义扭曲。
- 研究发现,尽管METEOR整合了语义特征,仍未能检测到情感反转,例如在否定被省略导致愤怒转为喜悦的案例中,METEOR得分为0.61。
- 结果表明,标准MT评估指标在评估UGC中情感保留方面能力不足,尤其在情绪强烈、非正式的文本(如推文)中表现不佳。
- 研究结论指出,当前自动评估方法在情感被扭曲时高估了翻译质量,呼吁在MT评估中引入情感感知指标以应对情感文本。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。