Skip to main content
QUICK REVIEW

[论文解读] Co-word Maps and Topic Modeling: A Comparison Using Small and Medium-Sized Corpora (n < 1000)

Loet Leydesdorff, Adina Nerghes|arXiv (Cornell University)|Nov 10, 2015
Topic Modeling参考文献 37被引用 7
一句话总结

本文比较了在中小型语料库(n < 1000)中使用共现词制图与主题建模的效果,发现两种方法产生的结果显著不相关。虽然共现词图能清晰识别语义成分,但主题模型揭示了非语义相似性(例如语言模式),表明在小型语料库中主题建模无法替代共现词制图,但在大规模语义制图中可能表现更优。

ABSTRACT

Induced by big data, has become an attractive alternative to mapping co-words in terms of co-occurrences and co-absences using network techniques. Does topic modeling provide an alternative for co-word mapping in research practices using moderately sized document collections? We return to the word/document matrix using first a single text with a strong argument (The Leiden Manifesto) and then upscale to a sample of moderate size (n = 687) to study the pros and cons of the two approaches in terms of the resulting possibilities for making semantic maps that can serve an argument. The results from co-word mapping (using two different routines) versus topic modeling are significantly uncorrelated. Whereas components in the co-word maps can easily be designated, the topic models provide sets of words that are very differently organized. In these samples, the topic models seem to reveal similarities other than semantic ones (e.g., linguistic ones). In other words, topic modeling does not replace co-word mapping in small and medium-sized sets; but the paper leaves open the possibility that topic modeling would work well for the semantic mapping of large sets.

研究动机与目标

  • 评估主题建模是否可作为中小型文档语料库中一种可行的共现词制图替代方法。
  • 评估在中等规模语料库(n = 687)中,主题模型的语义连贯性与可解释性是否优于共现词网络。
  • 探究主题模型在中小型文本语料库中是否揭示语义关系或其他潜在模式(如语言特征)。
  • 确定主题建模在何种条件下可能优于共现词制图用于语义制图。

提出的方法

  • 基于一篇强论证性文本(《莱顿宣言》)构建词/文档矩阵,以确立基线。
  • 扩展至687篇文档的样本,采用基于共现与共缺省的两种不同方法分析共现词制图。
  • 对同一语料库应用主题建模技术,生成文档中的主题分布。
  • 通过相关性分析比较共现词制图与主题建模生成的语义图结果。
  • 通过评估共现词图中的组件与主题模型中的主题是否可被有意义地标记,来评估可解释性。
  • 通过定性检查,探索主题模型输出中潜在的非语义驱动因素(如语言特征)。

实验结果

研究问题

  • RQ1主题建模在中小型语料库(n < 1000)中能否有效复制或替代共现词制图?
  • RQ2共现词制图识别出的语义成分与主题建模生成的主题在可解释性与连贯性方面如何比较?
  • RQ3主题模型中的词聚类主要由语义相似性驱动,还是由其他因素(如语言模式)主导?
  • RQ4在中等规模文档语料库中,共现词制图与主题建模的结果相关程度如何?
  • RQ5在何种条件下,主题建模可能比共现词制图更适合用于语义制图?

主要发现

  • 共现词制图与主题建模的结果显著不相关,表明两种方法生成的语义结构大相径庭。
  • 共现词图能够清晰且可解释地标识语义成分,而主题模型则将词语集合按非语义因素(如语言模式)组织。
  • 主题模型揭示的相似性并非主要源于语义,表明在中小型语料库中可能存在潜在的混淆因素。
  • 本研究未发现充分证据表明主题建模在中小型语料库中能可靠捕捉语义关系。
  • 尽管在小型语料库中存在局限性,本文仍保留可能性:主题建模在极大规模数据集中可能更适用于语义制图。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。