Skip to main content
QUICK REVIEW

[논문 리뷰] Content-based Text Categorization using Wikitology

Muhammad Rafi, Sundus Hassan|arXiv (Cornell University)|2012. 08. 17.
Text and Document Classification Technologies참고 문헌 14인용 수 15
한 줄 요약

이 논문은 위키토로지( Wikitology )라는 주제 맵 기반 지식 표현 시스템을 사용하여 콘텐츠 기반 텍스트 분류를 위한 의미적 유사도 측정 방법을 제안한다. 문서를 주제 맵으로 변환하고 공통 패턴 상관관계를 통해 유사도를 계산함으로써, 기준 데이터셋에서 기존의 벡터 기반 모델에 비해 텍스트 군집화 성능이 향상된다.

ABSTRACT

A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assign a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed semantic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns. The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.

연구 동기 및 목표

  • 문서 간 의미 관계를 포괄하지 못하는 기존의 벡터 기반 모델의 한계를 해결한다.
  • 문서 군집화에서 유사도 계산의 계산 부담을 완화한다.
  • 주제 맵을 의미적 유사도를 위한 구조화된 지식 표현 수단으로 탐색한다.
  • 외부 지식 소스에서 유래한 맥락적 및 의미적 관계를 통합함으로써 텍스트 군집화 성능을 향상시킨다.
  • 주제 맵 내 공통 패턴 기반의 새로운 유사도 측정 방법의 효과성을 입증한다.

제안 방법

  • 위키토로지 시스템을 활용해 문서를 주제 맵으로 변환한다. 이 시스템은 의미적 주제와 관계를 추출하고 정렬한다.
  • 엔티티와 그 관계를 명시적으로 모델링한 구조화된 지식 그래프로 각 문서를 표현한다.
  • 공통 패턴(반복되는 주제 구조) 간 상관관계를 통해 문서 간 유사도를 계산한다.
  • 주제 맵 기반 표현을 활용해 어휘적 중복을 초월한 의미적 관계를 포착한다.
  • 패턴 매칭과 구조적 유사도를 활용해 문서 간 의미적 일치 정도를 정량화한다.
  • 제안된 유사도 측정 방법을 표준 텍스트 마이닝 데이터셋에서 성능 평가를 위한 텍스트 군집화 파이프라인에 적용한다.

실험 결과

연구 질문

  • RQ1기존의 벡터 모델에 비해 주제 맵 기반 표현이 의미적 유사도 측정에서 텍스트 분류 성능을 향상시킬 수 있는가?
  • RQ2주제 맵 내 공통 패턴의 상관관계가 의미적으로 유의미한 방식으로 문서 유사도를 반영하는가?
  • RQ3제안된 방법이 문서 군집화에서 유사도 계산의 계산 부담을 어느 정도 감소시키는가?
  • RQ4지식 표현에 위키토로지를 사용할 경우 기준 데이터셋에서 군집화 정확도가 향상되는가?
  • RQ5의미적 유사도 측정 방법이 코사인 유사도나 워드넷 기반 접근 방식과 비교해 어떻게 성능을 내는가?

주요 결과

  • 제안된 주제 맵 기반 유사도 측정 방법은 기존의 전통적인 벡터 기반 모델에 비해 텍스트 군집화 성능에서 뛰어난 성능을 보였다.
  • 위키토로지에서 유래한 구조화된 지식을 활용해 문서 간 의미적 관계를 효과적으로 포착했다.
  • 텍스트 마이닝 데이터셋에서의 실험 결과, 새로운 유사도 측정 방법을 사용함으로써 군집화 정확도가 향상되었다.
  • 어휘 매칭에 의존하는 것에서 벗어나 의미 패턴과 구조적 상관관계에 중점을 두어 유사도 측정의 정교함을 높였다.
  • 주제 맵 내 공통 패턴의 상관관계가 문서 간 유사도의 강력한 지표로 입증되었다.
  • 연구 결과 주제 맵가 의미적 텍스트 분류를 위한 확장 가능하고 효과적인 프레임워크임을 확인했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.