[论文解读] A Challenge Set for French -> English Machine Translation.
本文提出一个包含506个句子的法语到英语机器翻译挑战集,旨在测试系统处理语言差异构造的能力,如形态句法、词句法、纯句法和纯词汇差异。在2017年10月和2018年1月对Google Translate和DeepL进行评估,DeepL在除形态句法外的所有类别中表现优于Google,整体成功率高出13%,在词汇和词句法类别中表现出可测量的进步。
We present a challenge set for French --> English machine translation based on the approach introduced in Isabelle, Cherry and Foster (EMNLP 2017). Such challenge sets are made up of sentences that are expected to be relatively difficult for machines to translate correctly because their most straightforward translations tend to be linguistically divergent. We present here a set of 506 manually constructed French sentences, 307 of which are targeted to the same kinds of structural divergences as in the paper mentioned above. The remaining 199 sentences are designed to test the ability of the systems to correctly translate difficult grammatical words such as prepositions. We report on the results of using this challenge set for testing two different systems, namely Google Translate and DEEPL, each on two different dates (October 2017 and January 2018). All the resulting data are made publicly available.
研究动机与目标
- 为法语→英语机器翻译开发一个针对特定语对的挑战集,以识别标准度量之外的持续性翻译难题。
- 评估最先进的神经机器翻译系统(Google Translate和DeepL)在特定语言差异问题上的表现。
- 提供一个公开可用的基准,用于追踪系统演进并识别在处理细微语法和句法差异方面仍存在的挑战。
提出的方法
- 人工构建了506个法语句子,专门针对特定的语言差异类型:形态句法、词句法、纯句法和纯词汇。
- 每个句子包含一个参考译文和一个是非问题,以聚焦评估特定语言问题。
- 评估协议采用人工判断来确定是非问题的正确性,两位作者解决分歧。
- 系统在两个时间点(2017年10月和2018年1月)进行测试,以追踪性能演变。
- 挑战集包含199个针对复杂语法词(如介词和代词)的示例,这些词需要上下文敏感的翻译。
- 数据和判断结果公开提供,以确保可复现性和可比较性。
实验结果
研究问题
- RQ1像Google Translate和DeepL这样的主流神经机器翻译系统在精心筛选的语言挑战性法语到英语翻译示例集上的表现如何?
- RQ2系统在处理形态句法差异(如前置代词、动词时态标记和格一致)方面的能力如何?
- RQ3系统在词汇和句法差异方面表现如何,特别是涉及介词和反身结构的情况?
- RQ4这些系统在2017年10月到2018年1月期间,在此挑战集上的表现有何演变?
主要发现
- DeepL在挑战集上的整体成功率比Google Translate高出13%,在纯词汇和词句法类别中差距最大。
- Google Translate从2017年10月到2018年1月整体表现提升了2.6%,而DeepL提升了1%,表明其进步较慢但稳定。
- 在形态句法差异方面,两个系统表现几乎相当,Google在同期提升了7%,而DeepL下降了4.6%。
- 在纯句法差异方面,Google Translate下降了3.5%,而DeepL的表现保持稳定,表明Google在句法处理方面出现退化。
- 挑战集显示,系统在介词和语法词方面仍存在显著困难,尤其当翻译选择依赖上下文时。
- 公开提供的数据集和评估判断使得对人工判断与系统判断的独立验证和比较成为可能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。