QUICK REVIEW

[论文解读] Domain, Translationese and Noise in Synthetic Data for Neural Machine Translation

Nikolay Bogoychev, Rico Sennrich|arXiv (Cornell University)|Nov 6, 2019

Natural Language Processing Techniques参考文献 45被引用 41

一句话总结

该论文将前向翻译和后向翻译作为神经机器翻译的数据增强进行分析，揭示测试数据的原始语言、翻译风格以及数据质量共同影响BLEU和人工评估；前向翻译在源语言为母语时通常有帮助，而后向翻译则能产生更流畅的输出。

ABSTRACT

The quality of neural machine translation can be improved by leveraging additional monolingual resources to create synthetic training data. Source-side monolingual data can be (forward-)translated into the target language for self-training; target-side monolingual data can be back-translated. It has been widely reported that back-translation delivers superior results, but could this be due to artefacts in the test sets? We perform a case study using French-English news translation task and separate test sets based on their original languages. We show that forward translation delivers superior gains in terms of BLEU on sentences that were originally in the source language, complementing previous studies which show large improvements with back-translation on sentences that were originally in the target language. To better understand when and why forward and back-translation are effective, we study the role of domains, translationese, and noise. While translationese effects are well known to influence MT evaluation, we also find evidence that news data from different languages shows subtle domain differences, which is another explanation for varying performance on different portions of the test set. We perform additional low-resource experiments which demonstrate that forward translation is more sensitive to the quality of the initial translation system than back-translation, and tends to perform worse in low-resource settings.

研究动机与目标

研究 synthetic data 来自 forward 和 backward translation 如何影响 NMT 性能。
考察测试集的原始语言（Original vs. Translationese）如何调节 synthetic data 的改进。
探索领域和翻译风格效应，以及 synthetic data 质量如何与 augmentation 方向相互作用。
评估 BLEU 是否与在 augmented MT 系统中的人工判断一致。

提出的方法

用 forward translation 和 back-translation 产生的合成数据训练双语 MT 系统（French→English），并与不使用合成数据的基线比较 transformer 与 RNN。
将测试集分成 Original（源语言为法语）与 Reverse（translationese French）部分，以评估方向特异的 BLEU。
在子集上进行人工评估，比较各系统的准确性与流畅性。
进行语言模型实验以区分翻译风格与领域效应。
用 Estonian→English 与 Finnish→English 测试泛化性，以分析数据质量敏感性。

实验结果

研究问题

RQ1Forward translation 是否在 Original 部分的测试数据上优于 back-translation，反之在 Reverse 部分？
RQ2翻译风格和领域差异如何解释 forward 与 back-translation 性能差异？
RQ3合成数据生成器的质量如何影响 forward 与 back-translation 的相对收益？
RQ4在自动 BLEU 分数与人工判断之间，是否对各种合成数据增强方法保持一致？
RQ5用合成数据训练的模型能否可靠地预测句子原始的源领域？

主要发现

系统	2008	2009	2010	2011	2012	2013
Original (French source) Baseline	29.4	44.2	32.9	32.3	37.3	47.4
Original (French source) BT_transformer	28.0	41.8	30.0	30.3	34.0	45.8
Original (French source) BT_rnn	29.3	42.2	31.7	31.5	34.6	46.9
Original (French source) FWD_transformer	29.0	43.8	32.3	32.4	36.4	49.0
Original (French source) FWD_rnn	30.9	45.1	32.0	33.1	38.3	48.3
Reverse (Translationese French source) Baseline	29.1	29.6	37.3	45.3	34.5	35.4
Reverse (Translationese French source) BT_transformer	31.6	32.9	42.6	50.8	39.3	39.5
Reverse (Translationese French source) BT_rnn	32.1	33.4	43.3	50.5	39.0	38.4
Reverse (Translationese French source) FWD_transformer	28.0	28.7	36.7	44.5	33.7	35.2
Reverse (Translationese French source) FWD_rnn	27.5	28.1	36.0	43.0	33.0	33.9
Full test set Baseline	29.2	37.3	35.2	38.8	35.9	41.6
Full test set BT_transformer	30.0	37.6	36.1	40.7	36.8	42.9
Full test set BT_rnn	30.9	38.1	37.3	40.5	36.9	42.5
Full test set FWD_transformer	28.5	36.7	34.4	38.5	35.0	42.2
Full test set FWD_rnn	29.0	37.1	33.9	38.0	35.6	41.5

Forward translation 常在 Original French source 部分的 BLEU 得分高于 back-translation。
Back-translation 通常在人工判断中提供更好的流畅度（在各个方向）。
BLEU 差异可能很大（多点量级），取决于测试集的原始语言，而人工可接受性差异较小。
语言模型分析表明翻译风格和领域差异都对观察到的效应有所贡献；在某些 EN↔FR 变体中效应达到平衡。
Forward translation 对初始翻译系统质量更敏感；在非常低质量的合成数据中，forward 的增益下降更多。

更好的研究，从现在开始

从论文设计到论文写作，大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成，并经人工编辑审核。