Skip to main content
QUICK REVIEW

[논문 리뷰] Artificial Intelligence Algorithms for Natural Language Processing and the Semantic Web Ontology Learning

Bryar A. Hassan, Tarik A. Rashid|arXiv (Cornell University)|2021. 08. 31.
Evolutionary Algorithms and Applications인용 수 12
한 줄 요약

이 논문은 이질적인 데이터셋에서 분할 정확도와 효율성을 향상시키는 개선된 진화적 군집화 알고리즘인 ECA*를 제안한다. ECA*는 온톨로지 학습에서 형식적 맥락 크기를 줄이는 데 적용되어 개념 래티스의 구조 품질을 89% 유지하면서 위키피디아 코퍼스에서의 개념 계층 추출을 가속화한다.

ABSTRACT

Evolutionary clustering algorithms have considered as the most popular and widely used evolutionary algorithms for minimising optimisation and practical problems in nearly all fields. In this thesis, a new evolutionary clustering algorithm star (ECA*) is proposed. Additionally, a number of experiments were conducted to evaluate ECA* against five state-of-the-art approaches. For this, 32 heterogeneous and multi-featured datasets were used to examine their performance using internal and external clustering measures, and to measure the sensitivity of their performance towards dataset features in the form of operational framework. The results indicate that ECA* overcomes its competitive techniques in terms of the ability to find the right clusters. Based on its superior performance, exploiting and adapting ECA* on the ontology learning had a vital possibility. In the process of deriving concept hierarchies from corpora, generating formal context may lead to a time-consuming process. Therefore, formal context size reduction results in removing uninterested and erroneous pairs, taking less time to extract the concept lattice and concept hierarchies accordingly. In this premise, this work aims to propose a framework to reduce the ambiguity of the formal context of the existing framework using an adaptive version of ECA*. In turn, an experiment was conducted by applying 385 sample corpora from Wikipedia on the two frameworks to examine the reduction of formal context size, which leads to yield concept lattice and concept hierarchy. The resulting lattice of formal context was evaluated to the original one using concept lattice-invariants. Accordingly, the homomorphic between the two lattices preserves the quality of resulting concept hierarchies by 89% in contrast to the basic ones, and the reduced concept lattice inherits the structural relation of the original one.

연구 동기 및 목표

  • 이질적인 데이터셋에서 기존 방법보다 뛰어난 분할 정확도와 강건성을 보이는 새로운 진화적 군집화 알고리즘 ECA*를 개발하는 것.
  • 특히 대규모 코퍼스에서 온톨로지 학습의 형식적 맥락 구축의 계산 비효율성을 해결하는 것.
  • 개념 래티스의 구조적 무결성을 훼손하지 않으면서 형식적 맥락의 모호성과 크기를 줄이는 것.
  • ECA*가 형식적 맥락을 최소화하면서도 개념 계층의 품질을 유지하는지 평가하는 것.
  • 세분화된 데이터 군집화를 통해 확장 가능하고 효율적인 온톨로지 학습을 가능하게 하는 것.

제안 방법

  • 다양하고 다기능적인 데이터셋에서 최적의 군집 성능을 달성하도록 설계된 적응형 진화적 군집화 알고리즘인 ECA*를 제안한다.
  • ECA*를 다섯 가지 최첨단 군집화 기법과 비교하기 위해 내부 및 외부 군집 평가 지표를 사용한다. 이 평가를 32개의 데이터셋에서 수행한다.
  • 의미적 코퍼스에서 관련 없거나 잘못된 속성-객체 쌍을 걸러내기 위해 ECA*를 적용하여 형식적 맥락 크기를 줄인다.
  • 원래 래티스와 감소된 래티스 간의 구조적 유사성을 비교하기 위해 개념 래티스 불변량을 사용한다.
  • 실제 NLP 및 온톨로지 학습 작업에서의 확장성과 성능 평가를 위해 385개의 위키피디아 코퍼스에 프레임워크를 적용한다.
  • 감소된 래티스가 원래 개념 래티스의 관계적 구조를 유지하고 있음을 확인하기 위해 동형 분석을 수행한다.

실험 결과

연구 질문

  • RQ1ECA*는 이질적인 데이터셋에서 기존의 진화적 군집화 알고리즘보다 분할 정확도와 안정성 측면에서 뛰어나게 작동할 수 있는가?
  • RQ2ECA* 기반의 형식적 맥락 감소는 원래 개념 래티스의 구조적 특성을 어느 정도 유지하는가?
  • RQ3ECA*는 온톨로지 학습에서 개념 래티스 생성의 계산 비용을 어느 정도 줄이는 데 효과적인가?
  • RQ4데이터셋의 특징이 ECA*의 군집화 및 맥락 감소 성능에 어떤 영향을 미치는가?
  • RQ5감소된 형식적 맥락은 후속 온톨로지 구축 작업에 충분한 품질을 유지하는가?

주요 결과

  • ECA*는 내부 및 외부 평가 지표를 모두 사용하여 32개의 이질적인 데이터셋에서 다섯 가지 최첨단 알고리즘과 비교해 뛰어난 군집 성능을 보였다.
  • ECA*를 형식적 맥락 감소에 적용한 결과, 래티스 불변량으로 측정한 바에 따르면 개념 래티스의 구조 품질을 89% 유지하였다.
  • 감소된 형식적 맥락은 의미적 정밀도를 훼손하지 않으면서도 개념 래티스 및 계층 추출에 대한 계산 시간을 크게 단축시켰다.
  • 원래 래티스와 감소된 래티스 간의 동형 관계는 개념 계층 내의 구조적 관계가 유지됨을 확인한다.
  • 적응형 ECA*는 노이즈가 많고 관련이 없는 속성-객체 쌍을 효과적으로 걸러내어 대규모 NLP 및 온톨로지 학습 파이프라인의 확장성을 향상시켰다.
  • 이 프레임워크는 385개의 위키피디아 코퍼스에 성공적으로 적용되어 실제 의미 데이터 처리에서 실용성과 성능 향상을 입증하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.