[论文解读] Multi-domain machine translation enhancements by parallel data extraction from comparable corpora
本文提出了一种无监督且可扩展的方法,从可比语料库(如维基百科或特定领域的网页资源集合)中提取高质量的平行句子,方法包括自动网页爬取和基于相似度的过滤。该方法通过从无现成平行资源的非平行来源中生成领域自适应的平行数据,显著提升了多种领域(如医学、电影、议会辩论)的统计机器翻译性能,BLEU得分实现可测量的提升。
Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from previously built comparable corpora. The methodologies are automatic and unsupervised which makes them good for large scale research. The task is highly practical as non-parallel multilingual data occur much more frequently than parallel corpora and accessing them is easy, although parallel sentences are a considerably more useful resource. In this study, we propose a method of automatic web crawling in order to build topic-aligned comparable corpora, e.g. based on the Wikipedia or Euronews.com. We also developed new methods of obtaining parallel sentences from comparable data and proposed methods of filtration of corpora capable of selecting inconsistent or only partially equivalent translations. Our methods are easily scalable to other languages. Evaluation of the quality of the created corpora was performed by analysing the impact of their use on statistical machine translation systems. Experiments were presented on the basis of the Polish-English language pair for texts from different domains, i.e. lectures, phrasebooks, film dialogues, European Parliament proceedings and texts contained medicines leaflets. We also tested a second method of creating parallel corpora based on data from comparable corpora which allows for automatically expanding the existing corpus of sentences about a given domain on the basis of analogies found between them. It does not require, therefore, having past parallel resources in order to train a classifier.
研究动机与目标
- 解决低资源和特定领域翻译任务中平行语料稀缺的问题。
- 开发一种无监督且可扩展的方法,从大规模可比语料库中挖掘平行句子。
- 通过从非平行源中提取的特定领域平行数据,丰富统计机器翻译系统,以实现领域自适应。
- 在无需预先存在平行训练数据的前提下,利用基于类比的句子匹配方法自动扩展现有平行语料库。
- 评估提取数据在多个领域和语言对上的翻译质量影响,特别是波兰语-英语语言对。
提出的方法
- 通过自动化网页爬取,从多语言来源(如维基百科)构建主题对齐的可比语料库。
- 应用无监督句子对齐技术,从可比语料库中识别候选平行句子。
- 使用基于相似度的过滤方法,去除不一致或部分等价的翻译,提升数据质量。
- 采用跨语言句子嵌入模型,检测跨语言的句子级对应关系。
- 设计一种方法,通过识别跨领域的类比句子结构,扩展现有平行语料库。
- 将提取的平行数据整合到统计机器翻译系统中,实现特定领域的适应。
实验结果
研究问题
- RQ1是否可以在无需预先存在平行数据的前提下,有效从可比语料库中提取平行句子?
- RQ2在翻译性能方面,所提取的平行数据质量与人工整理的语料库相比如何?
- RQ3该方法在医学、电影和议会听证会等多样化领域中,能在多大程度上提升机器翻译性能?
- RQ4该方法是否可推广至波兰语-英语以外的其他语言对?
- RQ5利用可比数据进行基于类比的现有平行语料库扩展,其有效性如何?
主要发现
- 所提出的方法成功从可比语料库中提取了高质量的平行句子,在多个领域显著提升了翻译性能。
- 使用提取的数据在所有测试领域(包括讲座、旅游手册、电影对白和医学说明书)中均带来了可测量的BLEU得分提升。
- 过滤机制有效降低了噪声,通过去除不一致或部分等价的翻译,提升了数据的可靠性。
- 基于类比的扩展方法在无需预先存在平行数据的前提下实现了语料库的扩展,展示了良好的可扩展性和适应性。
- 该方法在不同语言中均表现出有效性和可扩展性,在低资源和特定领域设置中观察到显著的性能提升。
- 评估结果证实,所提取的数据优于基于随机或未过滤可比数据训练的基线系统。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。