신효필 교수
Hyopil Shin
서울대학교 · 컴퓨터과학
연구실 소개
신효필 교수의 연구실은 한국어의 의미 구조와 언어 처리 기반의 지능형 정보 시스템 개발을 핵심으로 삼고 있습니다. 특히 한국어 어휘의 개념적 분류와 온톨로지 기반의 의미 모델링, 사용자 생성 텍스트의 특수성에 대응하는 정교한 토큰화 기법 개발에 주력하고 있습니다. 자연어 처리와 언어학 이론의 융합을 통해 한국어 특화의 정교한 언어 처리 기술을 구현하고자 합니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15In this paper I review two approaches - theoretical and statistical approach-, to Korean collocations and suggest a best statistical model for detecting collocation constructions. Pure theory-based approaches to collocations have problems of defining collocations and separating collocations from other free constructions or idioms. Statistical approaches, on the other hand, have been criticised because much focus lies only on frequencies of collocations and thus non-collocation constructions are
This paper describes the first year of work constructing the Korean Sentiment Corpus, focusing on the theoretical background such as the annotation scheme. Our aim is to provide a solid theoretical background for the corpus which reflects the characteristics of the Korean language and includes approximately 8,050 sentences taken from news articles. The corpus annotation scheme, based on the MPQA, is described along with the results of interannotator agreement tests with a view to improving the a
In this paper, I suggest a way of mapping the Mikrokosmos Ontology being developed at CRL (Computing Research Lab) of New Mexico State University, into lexical items. Many extensions to the Korean lexical meaning classifications have largely relied on the noun classifications and conceptual considerations. Those approaches, however, have proven to be insufficient in the case of Korean because lexical meaning classifications hardly fit in with conceptual structures. Along the same lines, simple m
User-generated texts include various types of stylistic properties, or noises. Such texts are not properly processed by existing morpheme analyzers or language models based on formal texts such as encyclopedias or news articles. In this paper, we propose a simple morphologically tight-fitting tokenizer (K-MT) that can better process proper nouns, coinages, and internet slang among other types of noise in Korean user-generated texts. We tested our tokenizer by performing classification tasks on K
In this paper, I suggested a way of separating concepts and actual words by mapping words to well-defined ontology. Other than previous work on the Korean wordnet, I claim that the separation words from concepts enables us to put syntactic and semantic constraints to the lexicon. Those constraints come from frame-based concepts in which various properties and relations are structured with slots. I chose 1283 Korean basic verbs presented by the National Academy of Korean Language and total 4818 s
We propose a method for dealing with semantic complexities occurring in information retrieval systems on the basis of linguistic observations. Our method follows from an analysis indicating that long runs of content words appear in a stopped document cluster, and our observation that these long runs predominately originate from the prepositional phrase and subject complement positions and as such, may be useful predictors of semantic coherence. From this linguistic basis, we test three statistic
This paper describes the two year endeavor of constructing the Korean Sentiment Analysis Corpus (KOSAC), focusing on the theoretical background and the analysis of the corpus itself. Our aim is to provide a solid theoretical background for the corpus which reflects the characteristics of the Korean language and includes approximately 7,744 sentences taken from news articles. The corpus annotation scheme, based on the MPQA, is described along with the statistics of features specified in the corpu
The LKB system is a grammar and lexicon development environment for use with constraint-based linguistic formalism. This system has led grammar engineers to achieve efficient processing and easy grammar development. We tried the flat Korean analysis in the LKB, motivated by relatively free word orders and vagueness of hierarchical VP structures in Korean, which further contributes to implementation difficulties. Since the LKB is basically designed for the binary analysis, our approach has fundam
대표 연구 분야
신효필 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.