Skip to main content
QUICK REVIEW

[Paper Review] Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks

Barack Wanjawa, Lilian Wanzare|arXiv (Cornell University)|Aug 25, 2022
Natural Language Processing Techniques4 citations
TL;DR

Kencorpus introduces a publicly available multilingual corpus of Swahili, Dholuo, and Luhya—three under-resourced Kenyan languages—featuring 5,594 items (5.6M words in text, 177 hours of speech). The dataset enables downstream NLP tasks such as POS tagging, machine translation, and question-answering, with proof-of-concept systems achieving 18.87% WER in speech-to-text and 80% EM in QA, significantly advancing low-resource language processing in East Africa.

ABSTRACT

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine learning and deep learning models without the requisite data. The Kencorpus project intends to bridge this gap by collecting and storing text and speech data that is good enough for data-driven solutions in applications such as machine translation, question answering and transcription in multilingual communities. The Kencorpus dataset is a text and speech corpus for three languages predominantly spoken in Kenya: Swahili, Dholuo and Luhya. Data collection was done by researchers from communities, schools, media, and publishers. The Kencorpus' dataset has a collection of 5,594 items - 4,442 texts (5.6M words) and 1,152 speech files (177hrs). Based on this data, Part of Speech tagging sets for Dholuo and Luhya (50,000 and 93,000 words respectively) were developed. We developed 7,537 Question-Answer pairs for Swahili and created a text translation set of 13,400 sentences from Dholuo and Luhya into Swahili. The datasets are useful for downstream machine learning tasks such as model training and translation. We also developed two proof of concept systems: for Kiswahili speech-to-text and machine learning system for Question Answering task, with results of 18.87% word error rate and 80% Exact Match (EM) respectively. These initial results give great promise to the usability of Kencorpus to the machine learning community. Kencorpus is one of few public domain corpora for these three low resource languages and forms a basis of learning and sharing experiences for similar works especially for low resource languages.

Motivation & Objective

  • To address the digital marginalization of indigenous African languages by creating a large-scale, publicly accessible corpus for under-resourced languages in Kenya.
  • To collect and curate high-quality text and speech data from local communities, schools, media, and publishers to support data-driven NLP applications.
  • To develop language-specific resources such as POS tags, question-answer pairs, and parallel translation sets for Swahili, Dholuo, and Luhya.
  • To demonstrate the feasibility of training and evaluating NLP models on low-resource African languages using the Kencorpus dataset.
  • To establish a foundation for future research and collaboration in low-resource language NLP, particularly in multilingual African contexts.

Proposed method

  • Data collection was conducted through community-based collaboration, sourcing text and speech from local schools, media outlets, publishers, and community members.
  • The corpus comprises 4,442 text items (5.6 million words) and 1,152 audio files (177 hours of speech) in Swahili, Dholuo, and Luhya.
  • Part-of-speech tagging was developed for Dholuo (50,000 words) and Luhya (93,000 words) using expert-annotated training data.
  • A question-answering dataset of 7,537 Swahili question-answer pairs was constructed from existing text sources.
  • A parallel translation set of 13,400 sentences was created, translating from Dholuo and Luhya into Swahili.
  • Two proof-of-concept systems were implemented: a speech-to-text system using automatic speech recognition and a question-answering model using supervised machine learning.

Experimental results

Research questions

  • RQ1Can a large-scale, community-driven corpus of under-resourced African languages like Swahili, Dholuo, and Luhya be effectively collected and curated for NLP applications?
  • RQ2To what extent can the Kencorpus dataset support downstream NLP tasks such as POS tagging, machine translation, and question-answering?
  • RQ3What performance can be achieved using the Kencorpus dataset in speech-to-text and question-answering systems for low-resource languages?
  • RQ4How does the quality and diversity of community-sourced data compare to traditional data collection methods in low-resource settings?
  • RQ5Can the Kencorpus framework be replicated for other low-resource African languages to improve digital inclusivity?

Key findings

  • The Kencorpus corpus contains 5,594 items, including 4,442 text documents (5.6 million words) and 1,152 audio files (177 hours of speech), representing a significant resource for three major Kenyan languages.
  • Part-of-speech tagging sets were successfully developed for Dholuo (50,000 words) and Luhya (93,000 words), enabling syntactic analysis in low-resource settings.
  • A question-answering dataset of 7,537 Swahili Q&A pairs was constructed, supporting training of extractive QA models.
  • A parallel translation set of 13,400 sentences was created, translating from Dholuo and Luhya into Swahili, facilitating multilingual NLP applications.
  • The speech-to-text system achieved a word error rate of 18.87%, demonstrating the feasibility of automatic speech recognition for these languages.
  • The question-answering model achieved an exact match score of 80%, indicating strong performance on extractive QA tasks using the Kencorpus data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.