Skip to main content
QUICK REVIEW

[논문 리뷰] A Semantic approach for effective document clustering using WordNet

Leena H. Patil, Mohammed Atique|arXiv (Cornell University)|2013. 03. 03.
Text and Document Classification Technologies참고 문헌 6인용 수 5
한 줄 요약

이 논문은 단어 네트워크(WordNet)를 사용하여 용어 선택을 향상시키고 차원을 감소시키는 의미적 문서 군집화 접근법을 제안한다. 동의어 및 의미 관계를 통합하여 적용하고, 전처리(불용어 제거, 어간 추출)를 수행하며, tf-idf, tf-df, tf2를 사용하여 용어 가중치를 설정함으로써, Reuters-21578 및 20 Newsgroups와 같은 벤치마크 데이터셋에서 군집화 정확도를 향상시킨다.

ABSTRACT

Now a days, the text document is spontaneously increasing over the internet, e-mail and web pages and they are stored in the electronic database format. To arrange and browse the document it becomes difficult. To overcome such problem the document preprocessing, term selection, attribute reduction and maintaining the relationship between the important terms using background knowledge, WordNet, becomes an important parameters in data mining. In these paper the different stages are formed, firstly the document preprocessing is done by removing stop words, stemming is performed using porter stemmer algorithm, word net thesaurus is applied for maintaining relationship between the important terms, global unique words, and frequent word sets get generated, Secondly, data matrix is formed, and thirdly terms are extracted from the documents by using term selection approaches tf-idf, tf-df, and tf2 based on their minimum threshold value. Further each and every document terms gets preprocessed, where the frequency of each term within the document is counted for representation. The purpose of this approach is to reduce the attributes and find the effective term selection method using WordNet for better clustering accuracy. Experiments are evaluated on Reuters Transcription Subsets, wheat, trade, money grain, and ship, Reuters 21578, Classic 30, 20 News group (atheism), 20 News group (Hardware), 20 News group (Computer Graphics) etc.

연구 동기 및 목표

  • 웹 및 이메일 시스템과 같은 대규모 텍스트 저장소에서 효과적인 문서 군집화 문제를 해결하기 위해.
  • 외부 지식으로서 WordNet을 활용하여 용어 간 의미 관계를 이용해 군집화 정확도를 향상시키기 위해.
  • 의미적 및 통계 기준에 기반한 지능적인 용어 선택을 통해 특징 공간의 차원을 감소시키기 위해.
  • 전통적인 용어 가중치 기법(tf-idf, tf-df, tf2)과 WordNet을 결합한 것이 문서 군집화에서 효과적인지 평가하기 위해.
  • 의미 인식 전처리 및 특징 선택을 통해 표준 텍스트 군집화 벤치마크에서 뛰어난 성능을 보여주기 위해.

제안 방법

  • 정규화를 위해 불용어 제거 및 Porter 어간 추출기를 사용하여 문서를 전처리한다.
  • WordNet을 활용하여 동의어 및 의미 관계를 식별하고, 전역 고유 단어 및 빈도 높은 단어 집합을 생성한다.
  • 전처리 및 의미적으로 풍부한 용어에서 구성된 용어-문서 행렬을 구축한다.
  • 최소 임계값을 사용하여 tf-idf, tf-df, tf2 등의 다중 용어 선택 방법을 적용한다.
  • 각 문서 내 용어 빈도를 세어 데이터 행렬에 표현한다.
  • WordNet의 의미 지식을 통합하여 용어 표현을 향상시키고 군집화 성능을 향상시킨다.

실험 결과

연구 질문

  • RQ1WordNet의 의미 정보를 통합할 경우, 기존 방법에 비해 문서 군집화 정확도가 어떻게 향상되는가?
  • RQ2의미적 풍부화와 결합했을 때, tf-idf, tf-df, tf2 중 어떤 용어 선택 전략이 가장 뛰어난 군집화 성능을 낼 수 있는가?
  • RQ3의미 전처리는 어떤 정도의 차원 감소를 이끌어내면서도 관련 문서 특징을 유지하는가?
  • RQ4Reuters-21578 및 20 Newsgroups와 같은 다양한 텍스트 컬렉션에서 제안된 방법은 어떻게 성능을 발휘하는가?
  • RQ5WordNet의 의미 관계는 군집화 작업에서 용어 표현을 효과적으로 향상시킬 수 있는가?

주요 결과

  • WordNet 통합은 의미 관계를 통해 용어 표현을 풍부화시켜 군집화 정확도를 크게 향상시킨다.
  • 의미 전처리와 결합했을 때, tf-df 및 tf2 방법이 tf-idf에 비해 군집화 품질에서 뛰어난 성능을 보였다.
  • 제안된 방법은 높은 관련성의 선택된 용어를 유지하면서 특징 공간을 효과적으로 감소시켰다.
  • Reuters-21578 및 20 Newsgroups 데이터셋에서의 실험 결과 군집화 F-측정치와 순수도에서 일관된 향상이 관찰되었다.
  • WordNet를 사용한 의미적 용어 풍부화로 전역 및 빈도 높은 단어 집합의 정확한 식별이 가능해져 전체 군집화 성능 향상에 기여했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.