[Paper Review] HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation
HOPE is a task-oriented, human-centric machine translation evaluation framework that uses professional post-editing annotations to assess MT quality with a focus on 'good enough' output. It employs a geometric progression scoring model for error penalty points across 8 key error types, achieving higher inter-rater reliability, faster evaluation, and strong alignment with human perception and post-editing effort—validated on English-Romanian marketing and business texts.
Traditional automatic evaluation metrics for machine translation have been widely criticized by linguists due to their low accuracy, lack of transparency, focus on language mechanics rather than semantics, and low agreement with human quality evaluation. Human evaluations in the form of MQM-like scorecards have always been carried out in real industry setting by both clients and translation service providers (TSPs). However, traditional human translation quality evaluations are costly to perform and go into great linguistic detail, raise issues as to inter-rater reliability (IRR) and are not designed to measure quality of worse than premium quality translations. In this work, we introduce HOPE, a task-oriented and human-centric evaluation framework for machine translation output based on professional post-editing annotations. It contains only a limited number of commonly occurring error types, and use a scoring model with geometric progression of error penalty points (EPPs) reflecting error severity level to each translation unit. The initial experimental work carried out on English-Russian language pair MT outputs on marketing content type of text from highly technical domain reveals that our evaluation framework is quite effective in reflecting the MT output quality regarding both overall system-level performance and segment-level transparency, and it increases the IRR for error type interpretation. The approach has several key advantages, such as ability to measure and compare less than perfect MT output from different systems, ability to indicate human perception of quality, immediate estimation of the labor effort required to bring MT output to premium quality, low-cost and faster application, as well as higher IRR. Our experimental data is available at \url{https://github.com/lHan87/HOPE}.
Motivation & Objective
- Address the limitations of traditional MT evaluation metrics, which lack semantic focus and fail to correlate with human judgment.
- Overcome the high cost, complexity, and low inter-rater reliability of standard human evaluation methods like MQM.
- Develop a scalable, task-oriented evaluation framework tailored for sub-premium, 'good enough' quality MT output in real-world industrial settings.
- Enable accurate estimation of post-editing effort and human perception of translation quality without tracking excessive linguistic details.
- Improve inter-rater reliability and reduce evaluation time while maintaining strong alignment with professional post-editing practices.
Proposed method
- Design a minimal set of 8 error types relevant to MT post-editing: proper name, impact, required adaptation, terminology, grammar, accuracy, style, and proofreading.
- Implement a geometric progression scoring model where error penalty points (EPPs) increase non-linearly with error severity, reflecting real post-editing effort.
- Apply the framework either during or after professional post-editing, using annotated translations as input for evaluation.
- Use segment-level scoring to enable transparency and system-level comparison across different MT engines.
- Train evaluators on a limited set of error categories to reduce learning curve and improve consistency.
- Validate the framework on two tasks: EN→RU marketing content (111 segments) and business domain text (3,339 words), using Google Translate and DeepL.
Experimental results
Research questions
- RQ1Can a minimal, error-type-focused evaluation framework achieve higher inter-rater reliability than traditional MQM-style methods?
- RQ2To what extent does the geometric progression scoring model reflect real post-editing effort and human perception of quality?
- RQ3How well does the HOPE framework correlate with human judgment for 'good enough' quality MT output, especially in technical and marketing domains?
- RQ4Can the framework be applied cost-effectively and quickly while still providing transparent, actionable feedback for MT system improvement?
- RQ5Does the framework effectively distinguish between different MT systems' performance on less-than-premium quality output?
Key findings
- The HOPE framework demonstrated high inter-rater reliability (IRR) in error type interpretation, significantly improving consistency over traditional methods.
- The geometric progression scoring model effectively reflected the severity of translation errors and correlated well with actual post-editing effort.
- Evaluation time was substantially reduced compared to full MQM-style assessments, while maintaining transparency at the segment level.
- The framework successfully measured and compared MT output quality across systems, even when output was below premium quality standards.
- The experimental results on English-Romanian marketing and business texts confirmed that HOPE aligns with human perception and provides actionable insights for MT improvement.
- The dataset and code are publicly available at https://github.com/lHan87/HOPE, enabling reproducibility and further research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.