Skip to main content
QUICK REVIEW

[논문 리뷰] Corpus-Based Approaches to Igbo Diacritic Restoration

Ignatius Ezeani|Lancaster EPrints (Lancaster University)|2026. 01. 26.
Natural Language Processing Techniques인용 수 1
한 줄 요약

이 박사학위 논문은 이보어의 다이아크리틱 복원에 대해 조사하고 세 가지 주요 접근 방식(표준 n-그램 모델, 분류 모델, 임베딩 모델)으로 유연한 데이터셋 생성 프레임워크를 제안한다.

ABSTRACT

With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in its syntax, pragmatics and phonology, which need to be captured and processed. The capacity of computers to process natural languages is increasing because NLP researchers are pushing its boundaries. But these research works focus more on well-resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese, etc. Over 95% of the world's 7000 languages are low-resourced for NLP, i.e. they have little or no data, tools, and techniques for NLP work. In this thesis, we present an overview of diacritic ambiguity and a review of previous diacritic disambiguation approaches on other languages. Focusing on the Igbo language, we report the steps taken to develop a flexible framework for generating datasets for diacritic restoration. Three main approaches, the standard n-gram model, the classification models and the embedding models were proposed. The standard n-gram models use a sequence of previous words to the target stripped word as key predictors of the correct variants. For the classification models, a window of words on both sides of the target stripped word was used. The embedding models compare the similarity scores of the combined context word embeddings and the embeddings of each of the candidate variant vectors.

연구 동기 및 목표

  • 저자원 언어를 위한 NLP의 필요성을 제고하고 이보어의 다이아크리틱 모호성을 해결한다.
  • 다양한 언어에서의 이전 다이아크리틱 구분(disambiguation) 접근법을 검토한다.
  • 이보어 다이아크리틱 복원을 위한 데이터셋 생성을 위한 유연한 프레임워크를 개발한다.

제안 방법

  • 이보어 다이아크리틱 복원을 위한 유연한 데이터셋 생성 프레임워크를 개발한다.
  • 세 가지 주요 모델링 접근 방식 제안: 표준 n-그램 모델, 분류 모델, 임베딩 모델.
  • 이보어의 다이아크리틱 예측을 위한 맥락 윈도우와 예측자 피처를 평가한다.

실험 결과

연구 질문

  • RQ1이보어의 다이아크리틱 모호성을 말뭉치 기반 접근법으로 어떻게 효과적으로 모델링할 수 있는가?
  • RQ2이보어 다이아크리틱 복원에 대한 n-그램, 분류, 임베딩 모델의 비교 이점은 무엇인가?
  • RQ3이보어 다이아크리틱 복원에 대해 유연한 평가를 가능하게 하는 데이터셋 생성 전략은 무엇인가?

주요 결과

  • 이보어 다이아크리틱 복원을 위한 세 가지 모델링 접근 방식이 제안된다: 표준 n-그램 모델, 분류 모델, 및 임베딩 모델.
  • 이보어에서 다이아크리틱 복원 연구를 지원하기 위한 데이터셋 생성 프레임워크가 개발된다.
  • 각 접근 방식은 올바른 다이아크리틱 변형을 예측하기 위해 대상에서 제거된 단어를 둘러싼 맥락을 사용한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.