[Paper Review] What Level of Quality can Neural Machine Translation Attain on Literary Text?
This study evaluates neural machine translation (NMT) on literary texts by training an NMT system and a phrase-based SMT (PBSMT) system on over 100 million words of parallel novel data. NMT significantly outperforms PBSMT, achieving a 3-point absolute BLEU improvement (11% relative) and producing translations of human-equivalent quality in 17–34% of cases, depending on the novel, according to human evaluation.
Given the rise of a new approach to MT, Neural MT (NMT), and its promising performance on different text types, we assess the translation quality it can attain on what is perceived to be the greatest challenge for MT: literary text. Specifically, we target novels, arguably the most popular type of literary text. We build a literary-adapted NMT system for the English-to-Catalan translation direction and evaluate it against a system pertaining to the previous dominant paradigm in MT: statistical phrase-based MT (PBSMT). To this end, for the first time we train MT systems, both NMT and PBSMT, on large amounts of literary text (over 100 million words) and evaluate them on a set of twelve widely known novels spanning from the the 1920s to the present day. According to the BLEU automatic evaluation metric, NMT is significantly better than PBSMT (p < 0.01) on all the novels considered. Overall, NMT results in a 11% relative improvement (3 points absolute) over PBSMT. A complementary human evaluation on three of the books shows that between 17% and 34% of the translations, depending on the book, produced by NMT (versus 8% and 20% with PBSMT) are perceived by native speakers of the target language to be of equivalent quality to translations produced by a professional human translator.
Motivation & Objective
- To assess the translation quality attainable by state-of-the-art neural machine translation (NMT) on literary texts, particularly novels.
- To compare NMT performance against the previous dominant paradigm, phrase-based statistical MT (PBSMT), on in-domain literary parallel data.
- To evaluate whether NMT can produce translations of quality comparable to professional human translations on literary content.
- To investigate the impact of textual features—lexical richness, novelty relative to training data, and sentence length—on NMT performance relative to PBSMT.
Proposed method
- Trained an NMT system and a PBSMT system on over 100 million words of parallel text from 12 widely known novels.
- Used automatic evaluation via the BLEU metric on all 12 novels to compare NMT and PBSMT performance.
- Conducted a human evaluation using the Appraise tool to rank translations from NMT, PBSMT, and professional human translators on three novels.
- Applied the TrueSkill algorithm to derive statistically significant overall scores for each translation type (HT, NMT, PBSMT) with p < 0.05.
- Analyzed the correlation between NMT’s relative improvement over PBSMT and three novel-specific features: lexical richness, novelty, and average sentence length.
- Used pairwise ranking analysis to compare the quality of NMT and PBSMT outputs across all three novels in the human evaluation.
Experimental results
Research questions
- RQ1Can NMT achieve significantly better translation quality than PBSMT on in-domain literary text, specifically novels?
- RQ2To what extent can NMT produce translations judged by native speakers to be of equivalent quality to professional human translations on literary content?
- RQ3How do textual characteristics—such as lexical richness, novelty relative to training data, and sentence length—affect the relative performance gain of NMT over PBSMT?
- RQ4What is the potential of NMT to assist professional literary translators through post-editing, based on quality and human perception?
Key findings
- NMT significantly outperformed PBSMT on all 12 novels according to the BLEU metric, with a 3-point absolute improvement (11% relative) and p < 0.01.
- In human evaluation, 17–34% of NMT translations were perceived as equivalent in quality to professional human translations, compared to 8–20% for PBSMT.
- For all three novels evaluated, NMT was ranked higher than PBSMT in 41.4–54.7% of cases, with PBSMT and NMT tied in 27.8–39.4% of cases.
- The overall human evaluation scores, derived via TrueSkill, showed that NMT ranked significantly above PBSMT and below human translations, with NMT achieving 18–22% of the distance toward human quality from the PBSMT baseline.
- Only sentence length showed a meaningful correlation with NMT’s relative improvement over PBSMT, with performance gains decreasing as sentence length increased.
- The novel with the longest average sentence length showed a relatively low improvement for NMT over PBSMT, suggesting sentence length as a key factor in performance variation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.