[Paper Review] Corpus-Based Approaches to Igbo Diacritic Restoration
This Ph.D. thesis surveys diacritic restoration for Igbo and proposes a flexible dataset generation framework with three main approaches: standard n-gram models, classification models, and embedding models.
With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in its syntax, pragmatics and phonology, which need to be captured and processed. The capacity of computers to process natural languages is increasing because NLP researchers are pushing its boundaries. But these research works focus more on well-resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese, etc. Over 95% of the world's 7000 languages are low-resourced for NLP, i.e. they have little or no data, tools, and techniques for NLP work. In this thesis, we present an overview of diacritic ambiguity and a review of previous diacritic disambiguation approaches on other languages. Focusing on the Igbo language, we report the steps taken to develop a flexible framework for generating datasets for diacritic restoration. Three main approaches, the standard n-gram model, the classification models and the embedding models were proposed. The standard n-gram models use a sequence of previous words to the target stripped word as key predictors of the correct variants. For the classification models, a window of words on both sides of the target stripped word was used. The embedding models compare the similarity scores of the combined context word embeddings and the embeddings of each of the candidate variant vectors.
Motivation & Objective
- Motivate NLP for low-resource languages and address diacritic ambiguity in Igbo.
- Review prior diacritic disambiguation approaches across languages.
- Develop a flexible framework to generate datasets for Igbo diacritic restoration.
Proposed method
- Develop a flexible dataset generation framework for Igbo diacritic restoration.
- Propose three main modeling approaches: standard n-gram models, classification models, embedding models.
- Evaluate context windows and predictor features for predicting diacritics in Igbo.
Experimental results
Research questions
- RQ1How can diacritic ambiguity in Igbo be effectively modeled using corpus-based approaches?
- RQ2What are the comparative advantages of n-gram, classification, and embedding models for Igbo diacritic restoration?
- RQ3What dataset generation strategies enable flexible evaluation for Igbo diacritic restoration?
Key findings
- Three modeling approaches are proposed for Igbo diacritic restoration: standard n-gram models, classification models, and embedding models.
- A dataset generation framework is developed to support diacritic restoration research in Igbo.
- Each approach uses context around the target stripped word to predict the correct diacritic variant.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.