Sungkyunkwan University · 情報科学
Professor Youngjoong Ko's research lab specializes in text mining, natural language processing, and information retrieval, with a strong focus on automatic text categorization and cross-language information retrieval. The lab explores unsupervised and semi-supervised learning techniques to reduce reliance on costly labeled data, emphasizing feature weighting, sentence importance estimation, and bootstrapping methods. It also develops language resources such as bilingual dictionaries and parallel corpora using large-scale web sources like Wikipedia to enhance multilingual text processing.
Figures are computed from collected data and may differ slightly.
No abstract available.
The goal of text categorization is to classify documents into a certain number of predefined categories. The previous works in this area have used a large number of labeled training documents for supervised learning. One problem is that it is difficult to create the labeled training documents. While it is easy to collect the unlabeled documents, it is not so easy to manually categorize them for creating training documents. In this paper, we propose an unsupervised learning method to overcome the
Automatic text categorization is a problem of automatically assigning text documents to predefined categories. In order to classify text documents, we must extract good features from them. In previous research, a text document is commonly represented by the term frequency and the inverted document frequency of each feature. Since there is a difference between important sentences and unimportant sentences in a document, the features from more important sentences should be considered more than oth
A wide range of supervised learning algorithms has been applied to Text Categorization. However, the supervised learning approaches have some problems. One of them is that they require a large, often prohibitive, number of labeled training documents for accurate learning. Generally, acquiring class labels for training data is costly, while gathering a large quantity of unlabeled data is cheap. We here propose a new automatic text categorization method for learning from only unlabeled data using
Open papers in the app to read, cite, and organize with AI.