Seung-Hoon Na
Korea University · Computer Science
About the Lab
Professor Seung-Hoon Na's research lab specializes in natural language processing and computational linguistics, with a focus on morphological analysis, term frequency normalization, and semantic representation in information retrieval. The lab develops advanced statistical and discriminative models—such as conditional random fields (CRFs)—for Korean morphological disambiguation and explores innovative methods to enhance document representation using translation-based term enrichment. Additionally, the lab investigates the role of semantic answer type taxonomies in improving question-answering systems and addresses fundamental challenges in text representation, including verbosity and topical scope in long documents. The research integrates linguistic theory with machine learning to build more robust and semantically aware NLP systems.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15A salient characteristic of sovereign defaults is that they are typically accompanied by large devaluations. This paper presents new evidence of this empirical regularity known as the Twin Ds and proposes a model that rationalizes it as an optimal policy outcome. The model combines limited enforcement of debt contracts and downward nominal wage rigidity. Under optimal policy, default is shown to occur during contractions. The role of default is to free up resources for domestic absorption, and t
There has been recent interest in statistical approaches to Korean morphological analysis. However, previous studies have been based mostly on generative models, including a hidden Markov model (HMM), without utilizing discriminative models such as a conditional random field (CRF). We present a two-stage discriminative approach based on CRFs for Korean morphological analysis. Similar to methods used for Chinese, we perform two disambiguation procedures based on CRFs: (1) morpheme segmentation an
Abstract. Term frequency normalization is a serious issue since lengths of doc-uments are various. Generally, documents become long due to two different rea-sons- verbosity and multi-topicality. First, verbosity means that the same topic is repeatedly mentioned by terms related to the topic, so that term frequency is more increased than the well-summarized one. Second, multi-topicality indicates that a document has a broad discussion of multi-topics, rather than single topic. Al-though these doc
The standard approach for term frequency normalization is based only on the document length. However, it does not distinguish the verbosity from the scope, these being the two main factors determining the document length. Because the verbosity and scope have largely different effects on the increase in term frequency, the standard approach can easily suffer from insufficient or excessive penalization depending on the specific type of long document. To overcome these problems, this article propos
In question answering (QA), answer types are semantic categories that questions require. An answer type taxonomy (ATT) is a collection of these answer types. ATT may heavily affect the performance of QA systems, because its broadness and granularity provides coverage and specificity of answer types. Cardie [1] used 13 categories for entity classification, and obtained large performance improvement, compared with the method using no categories. Also, according to Pasca et al. [3], the more catego
In this paper, we propose a novel phrase-based model for Korean morphological analysis by considering a phrase as the basic processing unit, which generalizes all the other existing processing units. The impetus for using phrases this way is largely motivated by the success of phrase-based statistical machine translation (SMT), which convincingly shows that the larger the processing unit, the better the performance. Experimental results using the SEJONG dataset show that the proposed phrase-base
Text retrieval queries frequently contain named entities. The standard approach of term frequency weighting does not work well when estimating the term frequency of a named entity, since anaphoric expressions (like he, she, the movie, etc) are frequently used to refer to named entities in a document, and the use of anaphoric expressions causes the term frequency of named entities to be underestimated. In this paper, we propose a novel 2-Poisson model to estimate the frequency of anaphoric expres
Research Areas
Dive deeper into Seung-Hoon Na's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.