[Paper Review] Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation
This paper introduces Hausa Visual Genome (HaVG), the first large-scale multimodal dataset for English-to-Hausa machine translation, comprising 32,923 image-caption pairs in both languages. The dataset is created by automatically translating English descriptions from the Hindi Visual Genome into Hausa, followed by careful post-editing with visual context, enabling improved multilingual NLP tasks for low-resource Hausa language applications.
Multi-modal Machine Translation (MMT) enables the use of visual information to enhance the quality of translations. The visual information can serve as a valuable piece of context information to decrease the ambiguity of input sentences. Despite the increasing popularity of such a technique, good and sizeable datasets are scarce, limiting the full extent of their potential. Hausa, a Chadic language, is a member of the Afro-Asiatic language family. It is estimated that about 100 to 150 million people speak the language, with more than 80 million indigenous speakers. This is more than any of the other Chadic languages. Despite a large number of speakers, the Hausa language is considered low-resource in natural language processing (NLP). This is due to the absence of sufficient resources to implement most NLP tasks. While some datasets exist, they are either scarce, machine-generated, or in the religious domain. Therefore, there is a need to create training and evaluation data for implementing machine learning tasks and bridging the research gap in the language. This work presents the Hausa Visual Genome (HaVG), a dataset that contains the description of an image or a section within the image in Hausa and its equivalent in English. To prepare the dataset, we started by translating the English description of the images in the Hindi Visual Genome (HVG) into Hausa automatically. Afterward, the synthetic Hausa data was carefully post-edited considering the respective images. The dataset comprises 32,923 images and their descriptions that are divided into training, development, test, and challenge test set. The Hausa Visual Genome is the first dataset of its kind and can be used for Hausa-English machine translation, multi-modal research, and image description, among various other natural language processing and generation tasks.
Motivation & Objective
- Address the scarcity of high-quality, diverse, and publicly available NLP resources for Hausa, a low-resource Afro-Asiatic language with over 100 million speakers.
- Overcome the limitations of existing Hausa NLP datasets, which are often scarce, machine-generated, or restricted to religious domains.
- Enable multimodal machine translation and image captioning for Hausa by providing a dataset that integrates visual context with linguistic descriptions.
- Bridge the research gap in low-resource language NLP by creating a benchmark dataset for Hausa-English translation and multimodal understanding.
- Support the development of robust, context-aware machine translation and vision-language models for Hausa, a language with high sociolinguistic relevance in West Africa.
Proposed method
- Adapt the Hindi Visual Genome (HVG) dataset by translating its English image descriptions into Hausa using an automatic machine translation system.
- Perform manual post-editing of the synthetic Hausa translations, using the corresponding images as context to correct errors and improve linguistic accuracy.
- Structure the final dataset into four splits: training (26,338), development (3,293), test (3,293), and challenge test (3,000) sets.
- Ensure linguistic and visual consistency by validating translations against image content, particularly to resolve ambiguities such as 'court' vs. 'tennis court'.
- Use the dataset to train and evaluate an image captioning model based on a single-layer LSTM decoder with cross-entropy loss and Adam optimization.
- Conduct both automatic (BLEU) and manual evaluation to assess caption quality, categorizing outputs by relevance to object of interest, region of interest, or incorrect content.
Experimental results
Research questions
- RQ1Can automatic translation of image descriptions from English to Hausa, followed by visual post-editing, produce a high-quality, contextually accurate multimodal dataset for low-resource languages?
- RQ2To what extent does visual context improve the accuracy of Hausa translations for ambiguous English phrases, such as 'court' or 'story'?
- RQ3How effective is the Hausa Visual Genome (HaVG) dataset in supporting image captioning and multilingual machine translation tasks compared to existing low-resource language benchmarks?
- RQ4What are the limitations of automatic evaluation metrics like BLEU in assessing image captioning performance for Hausa, and how do human-annotated categories compare?
- RQ5Can HaVG serve as a foundation for future shared tasks and extensions into other vision-language tasks such as Visual Question Answering (VQA) for Hausa?
Key findings
- The Hausa Visual Genome (HaVG) dataset contains 32,923 image-caption pairs in both English and Hausa, with balanced splits for training, development, test, and challenge evaluation.
- Manual post-editing using visual context significantly improved translation quality, resolving ambiguities such as 'court' (legal vs. sports) and 'story' (narrative vs. storey).
- Automatic evaluation using BLEU scores revealed low performance for image captioning models, indicating that n-gram overlap is insufficient for capturing semantic accuracy in Hausa.
- Manual evaluation showed that 68% of generated captions correctly described objects in the image, with 54% accurately describing the region of interest, highlighting the need for better evaluation metrics.
- The dataset is freely available under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license, enabling open research and future development in Hausa NLP.
- Future work includes creating a fully human-annotated version of HaVG without reliance on machine translation, and extending the dataset for Visual Question Answering (VQA) and shared tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.