Skip to main content
QUICK REVIEW

[论文解读] Moving Beyond LDA: A Comparison of Unsupervised Topic Modelling Techniques for Qualitative Data Analysis of Online Communities

Amandeep Kaur, James R. Wallace|arXiv (Cornell University)|Dec 19, 2024
Computational and Text Analysis Methods被引用 4
一句话总结

本研究评估了 BERTopic——一种基于大语言模型(LLM)的主题建模技术——在在线社区数据定性分析中相较于传统方法(LDA 和 NMF)的表现。研究发现,BERTopic 在生成连贯、详细且逻辑清晰的主题方面表现更优,能够揭示细微的关系;尽管在主题数量方面存在挑战,仍有 12 名参与者中的 8 名更倾向于使用 BERTopic 以获得更深层次的洞察与可操作发现。

ABSTRACT

Social media constitutes a rich and influential source of information for qualitative researchers. Although computational techniques like topic modelling assist with managing the volume and diversity of social media content, qualitative researcher's lack of programming expertise creates a significant barrier to their adoption. In this paper we explore how BERTopic, an advanced Large Language Model (LLM)-based topic modelling technique, can support qualitative data analysis of social media. We conducted interviews and hands-on evaluations in which qualitative researchers compared topics from three modelling techniques: LDA, NMF, and BERTopic. BERTopic was favoured by 8 of 12 participants for its ability to provide detailed, coherent clusters for deeper understanding and actionable insights. Participants also prioritised topic relevance, logical organisation, and the capacity to reveal unexpected relationships within the data. Our findings underscore the potential of LLM-based techniques for supporting qualitative analysis.

研究动机与目标

  • 评估 BERTopic——一种基于大语言模型(LLM)的主题建模技术——在在线社区定性数据分析中的可用性与有效性。
  • 解决定性研究者因缺乏编程技能而难以采用计算主题建模工具的障碍。
  • 从主题质量、可解释性及研究者偏好角度,对比 BERTopic 与传统方法(LDA 和 NMF)的表现。
  • 识别支持定性研究者分析大规模社交媒体数据的计算工具所需的关键设计需求。
  • 将 BERTopic 集成至计算主题分析(CTA)工具包中,并评估其对研究工作流程与洞察生成的影响。

提出的方法

  • 将 BERTopic 集成至 CTA 工具包,利用其基于变压器的词嵌入与上下文理解能力进行主题建模。
  • 将 BERTopic、LDA 和 NMF 应用于来自在线社区(特别是 Reddit)的定性数据集,以比较其主题输出结果。
  • 使用主题连贯性与多样性指标对模型性能进行定量评估,其中连贯性作为关键评估标准。
  • 对 12 名定性研究者开展半结构化访谈,以评估其对方法可用性、主题可解释性及偏好的看法。
  • 开展实际操作评估,让研究者在其自身数据集上应用 CTA 工具包,比较不同模型生成的主题簇。
  • 通过启用 GPU 加速并优化数据过滤与预处理流程,以满足 BERTopic 的计算需求。
Figure 1. Modeling and Sampling Module in the CTA Toolkit with BERTopic Integration: Model Details, a Topic List with keywords, Sample List of entries on the left side, and a Chord Graph for with Topics as Word clouds on the right. This layout facilitates detailed and comprehensive analysis of topic
Figure 1. Modeling and Sampling Module in the CTA Toolkit with BERTopic Integration: Model Details, a Topic List with keywords, Sample List of entries on the left side, and a Chord Graph for with Topics as Word clouds on the right. This layout facilitates detailed and comprehensive analysis of topic

实验结果

研究问题

  • RQ1与 LDA 和 NMF 相比,定性研究者如何看待 BERTopic 生成的主题在连贯性、相关性与可解释性方面的表现?
  • RQ2定性研究者在使用传统主题建模工具时面临的主要挑战是什么,特别是编程技能不足与数据预处理复杂性方面?
  • RQ3BERTopic 在何种方式上支持对在线社区定性数据中隐藏关系的深入理解与意外发现?
  • RQ4当选择主题建模方法时,研究者如何在主题组织、逻辑结构与可操作洞察之间进行优先级排序?
  • RQ5哪些设计特性——如可视化、搜索功能或层级结构——最能提升基于大语言模型的主题模型在定性研究中的可用性?

主要发现

  • 12 名定性研究者中有 8 名更偏好 BERTopic,因其能够生成详细、连贯且逻辑清晰的主题簇。
  • BERTopic 在主题连贯性与多样性指标上优于 LDA 和 NMF,表明其主题表征质量更高、更具意义。
  • 研究者高度认可 BERTopic 揭示数据中意外但重要关系的能力,从而实现更深入、更细致的分析。
  • 尽管 BERTopic 优势明显,但其生成的大量主题被部分参与者视为信息过载,凸显对更优层级可视化功能的需求。
  • LDA 和 NMF 因需处理无关符号与数学字符,需进行大量手动数据清洗,而 BERTopic 所需预处理较少。
  • CTA 工具包缺乏层级可视化被识别为一大局限,限制了研究者有效探索复杂主题结构的能力。
Figure 2. Chord diagram for LDA on the r/MachineLearning dataset, highlighting critiques of irrelevant symbols and the presence of mathematical symbols.
Figure 2. Chord diagram for LDA on the r/MachineLearning dataset, highlighting critiques of irrelevant symbols and the presence of mathematical symbols.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。