Skip to main content
QUICK REVIEW

[论文解读] A Continuous-time Mutually-Exciting Point Process Framework for Prioritizing Events in Social Media

Mehrdad Farajtabar, Safoora Yousefi|arXiv (Cornell University)|Nov 13, 2015
Human Mobility and Location-Based Analysis参考文献 28被引用 5
一句话总结

本文提出了一种特征调制的多变量 Hawkes 过程,通过建模自激、互激以及突发性互动动态,以优先排序社交媒体事件。通过整合文本、情感和时间特征,该方法在新闻动态优先排序任务中实现了最先进性能,其中语言学特征和用户资料特征对新事件排序最具预测力。

ABSTRACT

The overwhelming amount and rate of information update in online social media is making it increasingly difficult for users to allocate their attention to their topics of interest, thus there is a strong need for prioritizing news feeds. The attractiveness of a post to a user depends on many complex contextual and temporal features of the post. For instance, the contents of the post, the responsiveness of a third user, and the age of the post may all have impact. So far, these static and dynamic features has not been incorporated in a unified framework to tackle the post prioritization problem. In this paper, we propose a novel approach for prioritizing posts based on a feature modulated multi-dimensional point process. Our model is able to simultaneously capture textual and sentiment features, and temporal features such as self-excitation, mutual-excitation and bursty nature of social interaction. As an evaluation, we also curated a real-world conversational benchmark dataset crawled from Facebook. In our experiments, we demonstrate that our algorithm is able to achieve the-state-of-the-art performance in terms of analyzing, predicting, and prioritizing events. In terms of interpretability of our method, we observe that features indicating individual user profile and linguistic characteristics of the events work best for prediction and prioritization of new events.

研究动机与目标

  • 解决由于高流量、快速传播内容带来的社交媒体新闻动态优先排序挑战。
  • 在单一建模框架中统一静态(文本、情感)与动态(时间、基于互动)特征。
  • 通过基于特征的参数化方法,克服学习 Hawkes 过程参数时的数据稀缺问题。
  • 引入一种新颖的公开基准数据集,包含 50,000 条 Facebook 帖子和 100 万条评论,覆盖 16 个群组,用于评估优先排序算法。
  • 通过将特征权重与社交影响力及预测性能关联,提升模型可解释性。

提出的方法

  • 使用 Hawkes 过程将社交媒体互动建模为多变量点过程,以捕捉自激和互激现象。
  • 将强度函数参数化为特征(如语言学、用户资料、情感)的加权和,以减少参数数量并提升泛化能力。
  • 采用基于似然的学习目标,从观测到的互动序列中估计特征权重和激发核函数。
  • 通过时变非齐次强度函数引入时间动态,如突发性和衰减率。
  • 将模型应用于基于未来参与度预测结果,对用户新闻动态中的帖子进行优先排序。
  • 使用归一化平均排名(NAveRank)作为评估指标,以在不同用户群组和数据集之间比较性能。

实验结果

研究问题

  • RQ1如何统一社交媒体互动的静态与动态特征,以提升事件优先排序效果?
  • RQ2语言学特征和用户资料特征在多大程度上有助于预测社交媒体话题中的未来参与度?
  • RQ3特征调制的 Hawkes 过程是否能在新闻动态优先排序任务中超越基线模型?
  • RQ4数据稀疏性以及群组结构(如社交群组与主题群组)在多大程度上影响模型性能?
  • RQ5特征相关性与过拟合对模型泛化能力的影响如何,尤其是针对基于关系的特征?

主要发现

  • 所提出的特征调制 Hawkes 模型(HWK-ALL)在多样化 Facebook 群组中对社交媒体事件的优先排序任务中实现了最先进性能。
  • 语言学特征和用户资料特征始终优于基于关系的特征,后者因 Facebook 新闻动态排序中的数据偏差而容易过拟合。
  • 无特征 Hawkes 模型(HWK)在高级联量、用户数量较少的群组中表现良好,表明其对数据量和结构的敏感性。
  • 以社交关系为导向、由熟人驱动的互动群组(如本地居民、保龄球爱好者)比以主题为导向、一次性贡献的群组(如游戏玩家、马匹训练师)更具可预测性。
  • 在社交关联紧密的群组中,归一化平均排名(NAveRank)更低,证实了模型在这些场景下具有更高的预测能力。
  • 重叠或相关的特征(如基于关系的特征与语言学特征)即使在单个特征具有信息量时,仍会因噪声和过拟合而降低性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。