김길창 교수
Gil Chang Kim
KAIST 전산학부 · 컴퓨터과학
연구실 소개
김길창 교수의 연구실은 자연어 처리 및 기계 번역 분야에서 심층적인 연구를 수행하고 있습니다. 특히 한국어-영어 대화 번역을 위한 통계적 대화 분석 모델, 어순과 의미의 복잡성을 고려한 품사 태깅 기법, 그리고 문법적으로 어색한 문장을 자동으로 복구하는 강건한 구문 분석기 개발에 주력하고 있습니다. 연구는 어휘의 다의어 해소, 문맥 기반 번역 선택, 소규모 학습 데이터에서도 높은 성능을 내는 퍼지 기반 모델링 기법 등으로 이어져, 실제 응용에 유용한 기술적 기반을 구축하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15In some cases, to make a proper trans- lation of an utterance in a dialogue, the system needs various information about context. In this paper, we propose a sta- tistical dialogue analysis model based on speech acts for Korean-English dialogue machine translation. The model uses syn- tactic patterns and N-grams reflecting the hierarchical discourse structures of dia- logues. The syntactic pattern includes the syntactic features that are related with the language dependent expressions of speech a
Recently, most part-of-speech tagging approaches, such as rule-based, probabilistic and neural network approaches, have shown very promising results. In this paper, we are particularly interested in probabilistic approaches, which usually require lots of training data to get reliable probabilities. We alleviate such a restriction of probabilistic approaches by introducing a fuzzy network model to provide a method for estimating more reliable parameters of a model under a small amount of training
Part-of-speech (POS) tagging is a process of assigning a POS to each word in a sentence. Because many words are often ambiguous in their POSs, POS tagging must be able to select the most proper POS sequence for a given sentence. Recently, probabilistic approaches have shown very promising results to solve such ambiguity problems. Probabilistic approaches, however, usually require lots of training data to get reliable probabilities. To alleviate such restriction, we use fuzzy membership functions
An extragrammatical sentence is what a normal parser fails to analyze. It is important to recover it using only syntactic information although results of recovery are better if semantic factors are considered. A general algorithm for least-errors recognition, which is based only on syntactic information, was proposed by G. Lyon to deal with the extragrammaticality. We extended this algorithm to recover extragrammatical sentence into grammatical one in running text. Our robust parser with recover
In this paper, we introduce a method to represent phrase structure grammars for building a large annotated corpus of Korean syntactic trees. Korean is different from English in word order and word compositions. As a result of our study, it turned out that the differences are significant enough to induce meaningful changes in the tree annotation scheme for Korean with respect to the schemes for English. A tree annotation scheme defines the grammar formalism to be assumed, categories to be used, a
A word has many senses, and each sense can be mapped into many target words. Therefore, to select the appropriate translation with a correct sense, the sense of a source word should be disambiguated before selecting a target word. Based on this observation, we propose a hybrid method for translation selection that combines disambiguation of a source word sense and selection of a target word. Knowledge for translation selection is extracted from a bilingual dictionary and target language corpora.
In this paper, we propose a right-to-left dependency grammar parsing method for languages in which a governor appears after its modifier like Korean and Japanese. Unlike conventional left-to-right parsers, this parsing method can take advantage of the governor post-positioning property of such languages to reduce the size of search space by using the idea of a headable path. A headable path is a path which contains all candidate words which can be the governor of an input word during parsing. Th
In this paper, we propose a lexical selection method with three steps: sense disambiguation of source words, sense-to-word mapping, and selection of the most appropriate target language lexical item. The knowledge for each step is extracted from a machine readable dictionary and a target language monolingual corpus. By splitting the process of lexical selection into three steps and extracting the essential knowledge for each step from existing resources, our system can select appropriates word f
Translation selection is a process to select, from a set of target language words corresponding to a source language word, the most appropriate one that conveys the correct sense of a source word and makes the target language sentence more natural. In this paper, we propose a hybrid method for translation selection that exploits a bilingual dictionary and a target language corpus. Based on the ‘word-to-sense and sense-to-word ’ relationship between a source word and its translations, our method
대표 연구 분야
김길창 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.