고영중 교수
Young-Jin Ko
성균관대학교 소프트웨어학과 · 컴퓨터과학
연구실 소개
고영중 교수의 연구실은 텍스트 분류 및 정보 검색 분야에서 학습에 필요한 레이블이 적은 환경에서도 효과적으로 학습할 수 있는 비지도 학습 기반 텍스트 분류 기법을 주요 연구 주제로 다룹니다. 특히 문장의 중요도를 고려한 특성 가중치 부여, 키워드 기반 문장 분류, 그리고 부트스트래핑 기반의 레이블 없는 데이터 학습 기법을 통해 레이블링 비용을 줄이는 데 초점을 맞추고 있습니다. 또한, 다국어 정보 검색을 위한 번역 자원 구축과 병렬 코퍼스 생성 기법 개발을 통해 다국어 텍스트 처리 기술의 발전에도 기여하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15No abstract available.
The goal of text categorization is to classify documents into a certain number of predefined categories. The previous works in this area have used a large number of labeled training documents for supervised learning. One problem is that it is difficult to create the labeled training documents. While it is easy to collect the unlabeled documents, it is not so easy to manually categorize them for creating training documents. In this paper, we propose an unsupervised learning method to overcome the
Automatic text categorization is a problem of automatically assigning text documents to predefined categories. In order to classify text documents, we must extract good features from them. In previous research, a text document is commonly represented by the term frequency and the inverted document frequency of each feature. Since there is a difference between important sentences and unimportant sentences in a document, the features from more important sentences should be considered more than oth
A wide range of supervised learning algorithms has been applied to Text Categorization. However, the supervised learning approaches have some problems. One of them is that they require a large, often prohibitive, number of labeled training documents for accurate learning. Generally, acquiring class labels for training data is costly, while gathering a large quantity of unlabeled data is cheap. We here propose a new automatic text categorization method for learning from only unlabeled data using
대표 연구 분야
고영중 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.