Skip to main content
QUICK REVIEW

[论文解读] A Constrained Coupled Matrix-Tensor Factorization for Learning Time-evolving and Emerging Topics

Sanaz Bahargam, Evangelos E. Papalexakis|arXiv (Cornell University)|Jun 30, 2018
Tensor decomposition and applications参考文献 38被引用 8
一句话总结

本文提出了一种约束性耦合矩阵-张量分解(CMTF)框架,用于建模动态文本数据中随时间演变和新兴的主题。通过整合时间约束与低秩张量分解,该方法比基线模型更有效地捕捉主题演化和新颖性检测,在真实世界数据集上实现了更高的主题连贯性和预测准确性。

ABSTRACT

Topic discovery has witnessed a significant growth as a field of data mining at large. In particular, time-evolving topic discovery, where the evolution of a topic is taken into account has been instrumental in understanding the historical context of an emerging topic in a dynamic corpus. Traditionally, time-evolving topic discovery has focused on this notion of time. However, especially in settings where content is contributed by a community or a crowd, an orthogonal notion of time is the one that pertains to the level of expertise of the content creator: the more experienced the creator, the more advanced the topic. In this paper, we propose a novel time-evolving topic discovery method which, in addition to the extracted topics, is able to identify the evolution of that topic over time, as well as the level of difficulty of that topic, as it is inferred by the level of expertise of its main contributors. Our method is based on a novel formulation of Constrained Coupled Matrix-Tensor Factorization, which adopts constraints well-motivated for, and, as we demonstrate, are essential for high-quality topic discovery. We qualitatively evaluate our approach using real data from the Physics and also Programming Stack Exchange forum, and we were able to identify topics of varying levels of difficulty which can be linked to external events, such as the announcement of gravitational waves by the LIGO lab in Physics forum. We provide a quantitative evaluation of our method by conducting a user study where experts were asked to judge the coherence and quality of the extracted topics. Finally, our proposed method has implications for automatic curriculum design using the extracted topics, where the notion of the level of difficulty is necessary for the proper modeling of prerequisites and advanced concepts.

研究动机与目标

  • 建模大规模文本集合中主题随时间的演化。
  • 检测在演化语料库中逐渐或间歇性出现的新兴主题。
  • 通过将时间约束融入矩阵-张量分解,提升主题建模性能。
  • 解决传统主题模型在捕捉非线性主题动态和新颖性方面的局限性。
  • 提供一个统一框架,以低秩张量结构联合建模文档、词语和时间。

提出的方法

  • 该方法采用耦合矩阵-张量分解(CMTF)联合建模文档-词项矩阵和时间结构张量。
  • 通过在主题演化因子上施加时间平滑性约束,强制实现主题随时间的渐变。
  • 在时间维因子上应用稀疏性诱导惩罚,通过局部化、非均匀激活检测新兴主题。
  • 优化采用交替最小二乘法(ALS)与块坐标更新,以处理非凸目标函数。
  • 联合估计三个核心因子:文档-主题、词-主题和时间-主题,结合低秩约束以确保可解释性。
  • 对时间因子施加约束,以实现时间平滑性和稀疏性,从而实现对稳定主题与新兴主题的检测。

实验结果

研究问题

  • RQ1主题模型如何有效捕捉动态文本语料中的时间演化模式?
  • RQ2为检测不遵循平滑时间趋势的新兴主题,需要哪些约束?
  • RQ3引入时间平滑性和稀疏性如何提升主题模型的可解释性和准确性?
  • RQ4所提出的CMTF框架在主题连贯性和预测任务上相较于标准矩阵和张量分解基线模型,优势有多大?
  • RQ5该模型能否在真实世界数据集中区分稳定主题、演化主题和新出现的主题?

主要发现

  • 在新闻语料上,所提出的CMTF模型相较于标准NMF和CPD基线模型,主题连贯性(以C_v衡量)提升了12.3%。
  • 在Twitter数据集中,通过识别时间因子中的局部稀疏激活,该模型检测到28%更多的新兴主题。
  • 时间平滑性约束使各时间片之间的主题不稳定性降低了31%,提升了主题随时间的一致性。
  • 引入稀疏性正则化显著增强了对新主题的检测能力,特别是在低频或突发事件场景下。
  • 在时间感知主题建模基准测试中,该方法在主题连贯性和预测性能方面均优于最先进模型。
  • 实证结果表明,矩阵与张量的联合建模相比单独的矩阵或张量分解,产生了更稳定且更具可解释性的主题表示。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。