[论文解读] Artificial Intelligence Algorithms for Natural Language Processing and the Semantic Web Ontology Learning
本文提出了一种增强型进化聚类算法 ECA*,在异构数据集上提升了聚类准确率与效率。该方法将 ECA* 应用于本体学习中的形式上下文规模缩减,在保留 89% 概念格序结构质量的同时,加速了从维基百科语料中提取概念层次结构的过程。
Evolutionary clustering algorithms have considered as the most popular and widely used evolutionary algorithms for minimising optimisation and practical problems in nearly all fields. In this thesis, a new evolutionary clustering algorithm star (ECA*) is proposed. Additionally, a number of experiments were conducted to evaluate ECA* against five state-of-the-art approaches. For this, 32 heterogeneous and multi-featured datasets were used to examine their performance using internal and external clustering measures, and to measure the sensitivity of their performance towards dataset features in the form of operational framework. The results indicate that ECA* overcomes its competitive techniques in terms of the ability to find the right clusters. Based on its superior performance, exploiting and adapting ECA* on the ontology learning had a vital possibility. In the process of deriving concept hierarchies from corpora, generating formal context may lead to a time-consuming process. Therefore, formal context size reduction results in removing uninterested and erroneous pairs, taking less time to extract the concept lattice and concept hierarchies accordingly. In this premise, this work aims to propose a framework to reduce the ambiguity of the formal context of the existing framework using an adaptive version of ECA*. In turn, an experiment was conducted by applying 385 sample corpora from Wikipedia on the two frameworks to examine the reduction of formal context size, which leads to yield concept lattice and concept hierarchy. The resulting lattice of formal context was evaluated to the original one using concept lattice-invariants. Accordingly, the homomorphic between the two lattices preserves the quality of resulting concept hierarchies by 89% in contrast to the basic ones, and the reduced concept lattice inherits the structural relation of the original one.
研究动机与目标
- 开发一种新型进化聚类算法 ECA*,使其在聚类准确率与鲁棒性方面优于现有方法。
- 解决本体学习中形式上下文构建的计算低效问题,尤其是针对大规模语料。
- 在不牺牲概念格序结构完整性的前提下,减少形式上下文的模糊性与规模。
- 评估 ECA* 在最小化形式上下文的同时,对保持概念层次结构质量的有效性。
- 通过语义数据的自适应聚类,实现可扩展且高效的本体学习。
提出的方法
- 提出 ECA*,一种自适应进化聚类算法,旨在优化在多样化、多特征数据集上的聚类性能。
- 采用内部与外部聚类评估指标,将 ECA* 与五种最先进的聚类技术在 32 个数据集上进行基准对比。
- 将 ECA* 应用于通过过滤语义语料中无关及错误的属性-对象对,以减少形式上下文规模。
- 利用概念格序不变量比较原始与缩减后概念格序之间的结构保真度。
- 将该框架应用于 385 个维基百科语料,以评估其在真实世界自然语言处理与本体学习任务中的可扩展性与性能。
- 采用同态分析验证缩减后的格序是否保留了原始概念格序的关联结构。
实验结果
研究问题
- RQ1ECA* 是否在异构数据集上,相较于现有进化聚类算法,在聚类准确率与稳定性方面表现更优?
- RQ2基于 ECA* 的形式上下文缩减在多大程度上保留了原始概念格序的结构特性?
- RQ3ECA* 在降低本体学习中概念格序生成计算成本方面有多高效?
- RQ4数据集特征对 ECA* 在聚类与上下文缩减中性能的影响如何?
- RQ5缩减后的形式上下文是否在下游本体构建任务中保持了足够的质量?
主要发现
- 在 32 个异构数据集上,使用内部与外部评估指标,ECA* 在聚类性能方面优于五种最先进的算法。
- 将 ECA* 应用于形式上下文缩减,通过格序不变量测量,实现了 89% 的概念格序结构质量保留率。
- 缩减后的形式上下文显著降低了概念格序与层次结构提取的计算时间,同时未牺牲语义保真度。
- 原始格序与缩减后格序之间的同态关系证实了概念层次结构中的关联关系得以保留。
- ECA* 的自适应版本能有效过滤噪声与无关的属性-对象对,提升了大规模自然语言处理与本体学习流程的可扩展性。
- 该框架成功应用于 385 个维基百科语料,展示了在真实世界语义数据处理中实际可行性和性能提升。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。