Hyopil Shin
Seoul National University · Computer Science
About the Lab
Professor Hyopil Shin's research lab specializes in computational linguistics and natural language processing with a strong focus on the Korean language. The lab explores lexical semantics, ontology integration, and sentiment analysis, aiming to bridge theoretical linguistics with practical NLP applications. Key research directions include developing specialized corpora (e.g., Korean Sentiment Corpus), creating robust morphological tokenizers for user-generated content, and mapping lexical items to conceptual ontologies for improved semantic modeling.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15In this paper I review two approaches - theoretical and statistical approach-, to Korean collocations and suggest a best statistical model for detecting collocation constructions. Pure theory-based approaches to collocations have problems of defining collocations and separating collocations from other free constructions or idioms. Statistical approaches, on the other hand, have been criticised because much focus lies only on frequencies of collocations and thus non-collocation constructions are
This paper describes the first year of work constructing the Korean Sentiment Corpus, focusing on the theoretical background such as the annotation scheme. Our aim is to provide a solid theoretical background for the corpus which reflects the characteristics of the Korean language and includes approximately 8,050 sentences taken from news articles. The corpus annotation scheme, based on the MPQA, is described along with the results of interannotator agreement tests with a view to improving the a
In this paper, I suggest a way of mapping the Mikrokosmos Ontology being developed at CRL (Computing Research Lab) of New Mexico State University, into lexical items. Many extensions to the Korean lexical meaning classifications have largely relied on the noun classifications and conceptual considerations. Those approaches, however, have proven to be insufficient in the case of Korean because lexical meaning classifications hardly fit in with conceptual structures. Along the same lines, simple m
User-generated texts include various types of stylistic properties, or noises. Such texts are not properly processed by existing morpheme analyzers or language models based on formal texts such as encyclopedias or news articles. In this paper, we propose a simple morphologically tight-fitting tokenizer (K-MT) that can better process proper nouns, coinages, and internet slang among other types of noise in Korean user-generated texts. We tested our tokenizer by performing classification tasks on K
In this paper, I suggested a way of separating concepts and actual words by mapping words to well-defined ontology. Other than previous work on the Korean wordnet, I claim that the separation words from concepts enables us to put syntactic and semantic constraints to the lexicon. Those constraints come from frame-based concepts in which various properties and relations are structured with slots. I chose 1283 Korean basic verbs presented by the National Academy of Korean Language and total 4818 s
We propose a method for dealing with semantic complexities occurring in information retrieval systems on the basis of linguistic observations. Our method follows from an analysis indicating that long runs of content words appear in a stopped document cluster, and our observation that these long runs predominately originate from the prepositional phrase and subject complement positions and as such, may be useful predictors of semantic coherence. From this linguistic basis, we test three statistic
This paper describes the two year endeavor of constructing the Korean Sentiment Analysis Corpus (KOSAC), focusing on the theoretical background and the analysis of the corpus itself. Our aim is to provide a solid theoretical background for the corpus which reflects the characteristics of the Korean language and includes approximately 7,744 sentences taken from news articles. The corpus annotation scheme, based on the MPQA, is described along with the statistics of features specified in the corpu
The LKB system is a grammar and lexicon development environment for use with constraint-based linguistic formalism. This system has led grammar engineers to achieve efficient processing and easy grammar development. We tried the flat Korean analysis in the LKB, motivated by relatively free word orders and vagueness of hierarchical VP structures in Korean, which further contributes to implementation difficulties. Since the LKB is basically designed for the binary analysis, our approach has fundam
Research Areas
Dive deeper into Hyopil Shin's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.