Skip to main content
QUICK REVIEW

[论文解读] Retrofitting Concept Vector Representations of Medical Concepts to Improve Estimates of Semantic Similarity and Relatedness

Zhiguo Yu, Byron Wallace|PubMed|Sep 21, 2017
Biomedical Text Mining and Ontologies参考文献 13被引用 18
一句话总结

本文提出通过将分布式概念向量与UMLS元词汇表中的结构化知识进行融合,以改进生物医学领域语义相似性和相关性估计。通过将向量表示与UMLS关系对齐,该方法在UMNSRS基准测试中显著提升了与人类判断的相关性,优于以往的最先进方法。

ABSTRACT

Estimation of semantic similarity and relatedness between biomedical concepts has utility for many informatics applications. Automated methods fall into two categories: methods based on distributional statistics drawn from text corpora, and methods using the structure of existing knowledge resources. Methods in the former category disregard taxonomic structure, while those in the latter fail to consider semantically relevant empirical information. In this paper, we present a method that retrofits distributional context vector representations of biomedical concepts using structural information from the UMLS Metathesaurus, such that the similarity between vector representations of linked concepts is augmented. We evaluated it on the UMNSRS benchmark. Our results demonstrate that retrofitting of concept vector representations leads to better correlation with human raters for both similarity and relatedness, surpassing the best results reported to date. They also demonstrate a clear improvement in performance on this reference standard for retrofitted vector representations, as compared to those without retrofitting.

研究动机与目标

  • 解决分布式向量模型在捕捉生物医学概念分类结构方面的局限性。
  • 将UMLS元词汇表中的结构化知识整合到预训练的概念嵌入表示中。
  • 提升向量表示与人工标注的语义相似性和相关性判断的一致性。
  • 在标准的UMNSRS基准测试上评估微调在生物医学概念相似性中的有效性。
  • 证明结合分布式统计信息与知识图谱结构可提升语义表征质量。

提出的方法

  • 该方法利用UMLS元词汇表作为结构关系来源,对生物医学概念的预训练分布式上下文向量进行微调。
  • 采用加权平均过程,调整每个概念的向量,使其与元词汇表中语义关联的邻居向量更加相似。
  • 微调过程使用原始向量与邻居向量平均值的线性组合,其中包含可学习的权重参数。
  • 该方法在保留原始分布式语义的同时,融入了UMLS中的层次结构和关系结构。
  • 最终向量经过优化,以提升在人工标注基准上的相似性估计性能。
  • 该方法使用UMNSRS数据集进行评估,该数据集提供了生物医学概念的人工标注相似性和相关性评分。

实验结果

研究问题

  • RQ1使用UMLS元词汇表结构对分布式概念向量进行微调,能否提升语义相似性估计?
  • RQ2将生物医学知识图谱中的结构化知识融入,是否能超越仅依赖分布式模型的相关性估计性能?
  • RQ3微调向量在UMNSRS基准测试中的性能与最先进方法相比如何?
  • RQ4整合词汇与结构信息在多大程度上提升了与人类判断的一致性?
  • RQ5使用微调向量时,与人类评分者之间的相关性是否可测量地提升?

主要发现

  • 该微调方法在UMNSRS基准测试中与人类评分者的相关性高于以往报告的所有方法。
  • 该方法显著提升了语义相似性和相关性估计性能,超越了迄今为止报告的最佳结果。
  • 在标准的UMNSRS评估集上,微调向量相比非微调向量表现出明显的性能提升。
  • 该方法有效将UMLS的结构化知识整合到分布式向量中,同时未丢弃经验分布式统计信息。
  • 结果表明,结合分布式信息与结构化信息可生成更准确、更符合人类判断的生物医学语义表征。
  • 该改进在相似性与相关性任务中均保持一致,表明该方法具有广泛的适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。