[Paper Review] Considerations for meaningful sign language machine translation based on glosses
This paper critiques current practices in sign language machine translation using glosses, identifying critical flaws in dataset transparency, evaluation standards, and baseline rigor. It advocates for corpus-specific gloss processing, standardized evaluation via SacreBLEU, stronger baselines, and explicit discussion of gloss limitations to improve the credibility and impact of future research.
Automatic sign language processing is gaining popularity in Natural Language Processing (NLP) research (Yin et al., 2021). In machine translation (MT) in particular, sign language translation based on glosses is a prominent approach. In this paper, we review recent works on neural gloss translation. We find that limitations of glosses in general and limitations of specific datasets are not discussed in a transparent manner and that there is no common standard for evaluation. To address these issues, we put forward concrete recommendations for future research on gloss translation. Our suggestions advocate awareness of the inherent limitations of gloss-based approaches, realistic datasets, stronger baselines and convincing evaluation.
Motivation & Objective
- To identify and highlight the under-discussed limitations of gloss-based sign language translation in recent NLP research.
- To address the lack of standardized evaluation practices, especially regarding BLEU scoring and metric consistency.
- To encourage the use of realistic, diverse datasets beyond the PHOENIX corpus to improve generalizability.
- To promote corpus-specific gloss preprocessing informed by transcription conventions.
- To advocate for optimized baselines and reproducible code to enhance research credibility and reproducibility.
Proposed method
- Conducted a systematic review of 14 recent gloss translation papers to assess their treatment of gloss limitations, dataset use, evaluation methods, and baseline strength.
- Evaluated the consistency and transparency of BLEU computation, recommending the use of SacreBLEU with disabled internal tokenization for fairness.
- Proposed corpus-specific gloss preprocessing to account for variations in gloss transcription conventions across datasets.
- Recommended optimization of baselines using established low-resource MT techniques such as label smoothing and subword vocabulary tuning.
- Advocated for the inclusion of limitations in research papers, particularly in ACL-style limitation sections, to improve methodological transparency.
- Stressed the importance of reproducible code and open sharing of training procedures to support validation and extension of results.
Experimental results
Research questions
- RQ1To what extent are the inherent limitations of gloss-based sign language translation acknowledged in recent research?
- RQ2How consistent and reliable are evaluation practices—particularly BLEU scoring—across gloss translation studies?
- RQ3What are the consequences of relying on narrow, domain-specific datasets like PHOENIX for generalization in sign language MT?
- RQ4How can gloss preprocessing be adapted to reflect corpus-specific transcription conventions without introducing bias?
- RQ5Is gloss-based translation still the most effective approach, or should research prioritize linguistic tools like segmentation or coreference resolution?
Key findings
- Eight out of fourteen reviewed papers failed to adequately discuss the limitations of gloss-based approaches, leading to an overstatement of their practical utility.
- There is no standardized evaluation method across gloss translation papers, with inconsistent BLEU computation practices and lack of metric signatures.
- The PHOENIX dataset is frequently used but is limited in size and linguistic domain, yet its constraints are rarely acknowledged.
- Glosses are corpus-specific and vary significantly in transcription conventions, making cross-corpus comparisons invalid without careful preprocessing.
- Weak and unoptimized baselines are commonly used, undermining the validity of reported improvements in model performance.
- The use of SacreBLEU with disabled internal tokenization is recommended to ensure fair and comparable BLEU scores across studies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.