Skip to main content
QUICK REVIEW

[Paper Review] NLP for Ghanaian Languages

Paul Azunre, Salomey Osei|arXiv (Cornell University)|Mar 29, 2021
Topic Modeling14 references4 citations
TL;DR

NLP Ghana proposes an open-source initiative to develop state-of-the-art natural language processing tools for low-resource Ghanaian languages, focusing on Akan, Ewe, and Ga. By leveraging crowd-sourced and human-verified translation data, the team has created over 25,000 high-quality sentence pairs in Akuapem Twi and open-sourced models and datasets to advance machine translation, NER, and POS tagging for under-resourced African languages.

ABSTRACT

NLP Ghana is an open-source non-profit organization aiming to advance the development and adoption of state-of-the-art NLP techniques and digital language tools to Ghanaian languages and problems. In this paper, we first present the motivation and necessity for the efforts of the organization; by introducing some popular Ghanaian languages while presenting the state of NLP in Ghana. We then present the NLP Ghana organization and outline its aims, scope of work, some of the methods employed and contributions made thus far in the NLP community in Ghana.

Motivation & Objective

  • Address the lack of NLP tools and training data for Ghanaian languages, which are underrepresented in digital and machine learning systems.
  • Overcome the scarcity of annotated, high-quality training data for low-resource African languages like Akan, Ewe, and Ga.
  • Develop open-source, reusable NLP models and datasets to support machine translation, named entity recognition, and part-of-speech tagging.
  • Establish a sustainable, community-driven research organization to advance NLP for Ghanaian and other low-resource languages.
  • Improve the digital presence and computational usability of Ghanaian languages to support education, health services, and national security applications.

Proposed method

  • Employed crowd-sourcing via Google Forms to collect voluntary English-to-Akan translations, generating approximately 697 sentence pairs.
  • Used human-verification of machine-translated outputs to produce ~25,000 high-quality Akuapem Twi sentence pairs from an initial pool of ~50,000 machine translations.
  • Leveraged external multilingual datasets such as JW300 and the Bible to augment internal data collection efforts.
  • Developed and open-sourced contextual (BERT-based) and static (fastText) embeddings for Akan, Ewe, and Ga using the collected training data.
  • Organized a multi-team structure—Data, Engineering, Research, and Communications—to coordinate data collection, model development, and stakeholder engagement.
  • Planned future expansion to include audio data (oral corpora) and annotation of datasets for downstream tasks like NER and POS tagging.

Experimental results

Research questions

  • RQ1How can high-quality, low-resource training data be collected for Ghanaian languages like Akuapem Twi with limited funding and professional translator access?
  • RQ2To what extent can crowd-sourced and human-verified machine translation data produce reliable, usable NLP models for under-resourced African languages?
  • RQ3Can open-source, community-driven NLP initiatives effectively develop and deploy neural models (e.g., BERT, fastText) for low-resource languages like Akan and Ewe?
  • RQ4What are the key challenges in building NLP infrastructure for tonal, multilingual, and low-resource languages in Ghana, and how can they be mitigated?
  • RQ5How can NLP tools for Ghanaian languages contribute to national priorities such as education, health, and cybersecurity?

Key findings

  • NLP Ghana successfully collected and verified approximately 25,000 high-quality English-to-Akuapem Twi sentence pairs through a hybrid approach of machine translation and human review.
  • The organization has open-sourced BERT-based and fastText embeddings for Akan, Ewe, and Ga, enabling downstream NLP tasks such as classification and named entity recognition.
  • Over 697 sentence pairs were collected via crowd-sourcing, demonstrating the viability of community-driven data collection for low-resource languages.
  • The initiative has built a multi-disciplinary, open-source research community with over 100 members across academia, industry, and local language expertise.
  • Despite financial constraints, the team achieved significant progress in data collection and model development, laying the foundation for scalable NLP infrastructure in Ghana.
  • The project highlights the urgent need for funding and institutional support to expand data collection to audio and other modalities, and to improve annotation quality for advanced NLP tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.