[논문 리뷰] Probabilistic Modelling of Morphologically Rich Languages
이 학위 논문은 문법적으로 풍부한 언어에서 형태소 하위 구조를 명시적으로 통합하는 확률적 언어 모델을 제안한다. 베이지안 n-그램 모델과 분산 형태소 표현을 사용하여 스무딩과 일반화를 향상시킨다. 이는 연속적이고 비연속적인 형태소를 동시에 탐지할 수 있는 새로운 모델을 도입하여 형태소 분할 및 기계 번역과 같은 후속 작업에서 향상된 성능을 입증한다.
This thesis investigates how the sub-structure of words can be accounted for in probabilistic models of language. Such models play an important role in natural language processing tasks such as translation or speech recognition, but often rely on the simplistic assumption that words are opaque symbols. This assumption does not fit morphologically complex language well, where words can have rich internal structure and sub-word elements are shared across distinct word forms. Our approach is to encode basic notions of morphology into the assumptions of three different types of language models, with the intention that leveraging shared sub-word structure can improve model performance and help overcome data sparsity that arises from morphological processes. In the context of n-gram language modelling, we formulate a new Bayesian model that relies on the decomposition of compound words to attain better smoothing, and we develop a new distributed language model that learns vector representations of morphemes and leverages them to link together morphologically related words. In both cases, we show that accounting for word sub-structure improves the models' intrinsic performance and provides benefits when applied to other tasks, including machine translation. We then shift the focus beyond the modelling of word sequences and consider models that automatically learn <em>what</em> the sub-word elements of a given language are, given an unannotated list of words. We formulate a novel model that can learn discontiguous morphemes in addition to the more conventional contiguous morphemes that most previous models are limited to. This approach is demonstrated on Semitic languages, and we find that modelling discontiguous sub-word structures leads to improvements in the task of segmenting words into their contiguous morphemes.
연구 동기 및 목표
- 서브워드 구조를 활용하여 형태적으로 풍부한 언어의 언어 모델링에서 데이터 희소성 문제를 해결한다.
- 복합어의 베이지안 분해를 통해 n-그램 모델의 스무딩과 일반화를 향상시킨다.
- 형태소의 벡터 표현을 학습하는 분산 언어 모델을 개발하여 관련된 단어 형태 간의 유사성을 연결한다.
- 정렬되지 않은 단어 목록에서 연속형과 비연속형 형태소를 포함한 형태적 단위를 비지도로 탐지할 수 있도록 한다.
- 서브워드 모델링이 형태소 분할 및 기계 번역 작업에서 어떻게 유용한지 입증한다.
제안 방법
- 복합어를 형태소로 분해하여 스무딩을 향상시키고 데이터 희소성을 다루는 베이지안 n-그램 모델을 제안한다.
- 형태소의 조밀한 벡터 표현을 학습하고 이를 단어 형태 유사성 모델링에 활용하는 분산 언어 모델을 설계한다.
- 비연속 형태소를 명시적으로 다룰 수 있는 새로운 확률적 모델을 제안하여 비지도 형태소 탐지 문제를 해결한다.
- 형태소 복잡도가 높고 비연속 형태소가 흔한 세미틱 언어에 모델을 적용한다.
- 모델 파라미터 추정 및 형태소 구조 추론을 위해 변분 추론과 비모수 베이지안 방법을 사용한다.
- 학습된 형태소 표현을 기계 번역 및 형태소 분할과 같은 후속 자연어 처리 작업에 통합한다.
실험 결과
연구 질문
- RQ1단어 서브구조를 모델링하면 내재된 언어 모델 성능과 데이터 효율성이 향상되는가?
- RQ2분산 형태소 표현이 형태적으로 관련된 단어들 간의 일반화 능력을 얼마나 향상시키는가?
- RQ3확률적 모델이 형태적으로 복잡한 언어에서 정렬되지 않은 단어 목록에서 비연속 형태소를 탐지할 수 있는가?
- RQ4서브구조 모델링이 형태소 분할 및 기계 번역 성능에 어떤 영향을 미치는가?
- RQ5형태소 분해를 통합하면 n-그램 모델에서 스무딩이 향상되고 퍼즐리티가 낮아지는가?
주요 결과
- 형태소 분해를 통한 베이지안 n-그램 모델은 표준 n-그램 모델 대비 더 나은 스무딩과 낮은 퍼즐리티를 달성한다.
- 학습된 형태소 벡터를 활용한 분산 언어 모델은 일반화 능력 향상과 함께 형태적으로 풍부한 테스트 셋에서 낮은 퍼즐리티를 기록한다.
- 제안된 비지도 모델은 세미틱 언어에서 연속형과 비연속형 형태소를 성공적으로 식별하며, 연속형 단위에 국한된 모델보다 뛰어난 성능을 보인다.
- 발견된 형태소 구조를 사전 또는 특징 입력으로 사용할 경우 형태소 분할 성능이 향상된다.
- 형태소 기반 언어 모델을 통합한 신경 기계 번역 시스템에서 측정 가능한 성능 향상이 관찰된다.
- 특히 축합어 및 융합어 언어에서 다양한 형태학적 패러다임에 대해 강건한 성능을 보여준다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.