[论文解读] Expecting to be HIP: Hawkes Intensity Processes for Social Media Popularity
本文提出霍克斯强度过程(HIP),一种新颖的数学模型,用于量化外部推广(如推文)如何驱动YouTube等平台上的视频流行度。通过建模外部刺激与内在病毒式传播之间的相互作用,HIP在仅基于历史数据的基线模型上将流行度预测准确率提高了28.6%,并能识别出具有高病毒传播潜力或对推广不敏感的视频。
Modeling and predicting the popularity of online content is a significant problem for the practice of information dissemination, advertising, and consumption. Recent work analyzing massive datasets advances our understanding of popularity, but one major gap remains: To precisely quantify the relationship between the popularity of an online item and the external promotions it receives. This work supplies the missing link between exogenous inputs from public social media platforms, such as Twitter, and endogenous responses within the content platform, such as YouTube. We develop a novel mathematical model, the Hawkes intensity process, which can explain the complex popularity history of each video according to its type of content, network of diffusion, and sensitivity to promotion. Our model supplies a prototypical description of videos, called an endo-exo map. This map explains popularity as the result of an extrinsic factor - the amount of promotions from the outside world that the video receives, acting upon two intrinsic factors - sensitivity to promotion, and inherent virality. We use this model to forecast future popularity given promotions on a large 5-months feed of the most-tweeted videos, and found it to lower the average error by 28.6% from approaches based on popularity history. Finally, we can identify videos that have a high potential to become viral, as well as those for which promotions will have hardly any effect.
研究动机与目标
- 建模在持续外部影响下的在线内容流行度的复杂动态。
- 量化外部推广(如推特上的推广)与内在流行度增长(如YouTube上的增长)之间的关系。
- 开发一个结合外部刺激与内在病毒传播潜力的预测模型。
- 识别具有高病毒传播潜力或对推广不敏感的视频。
提出的方法
- 提出霍克斯强度过程(HIP),即霍克斯点过程的扩展,用于建模预期事件数量而非事件发生时间。
- 通过随机事件历史取期望,推导出用于流行度预测的确定性强度函数。
- 引入两个关键指标:内生响应(固有的病毒传播性)和外生敏感性(对推广的响应)。
- 采用内生-外生映射(endo-exo map),一种二维可视化工具,结合敏感性与内生响应,对视频潜力进行分类。
- 将HIP应用于包含8190万条与10.6亿条推文关联的5个月YouTube视频数据集,使用#shares和#tweets作为外生输入。
- 使用配对T检验和两样本T检验,验证HIP在预测性能上相对于基线MLR模型的表现。
实验结果
研究问题
- RQ1持续的外部推广在多大程度上影响在线内容的流行度动态?
- RQ2在已知计划的外部推广下,能否预测视频未来的流行度?
- RQ3哪些视频最有可能成为病毒式传播内容,哪些对推广不敏感?
- RQ4不同外部刺激(如#shares与#tweets)在预测能力上如何比较?
- RQ5能否使用统一的数学框架建模内在病毒传播性与外部推广之间的相互作用?
主要发现
- 与基于流行度历史的模型(如MLR)相比,HIP将平均预测误差降低了28.6%。
- 在经历强烈外生冲击的视频上,该模型实现了3.25%的中位数绝对预测误差,显著优于MLR的6.5%中位数误差。
- 使用#shares或#tweets作为外部刺激之间未发现统计学上显著差异,两者表现近乎一致(HIP的Cohen’s d = -0.05,MLR的Cohen’s d = 0.00)。
- #shares与#tweets时间序列之间的相关性在各类视频中均较高,皮尔逊相关系数的均值为0.78,中位数为0.87。
- HIP在预测性能上显著优于MLR,p值小于10−95,效应量(Cohen’s d)为0.197至0.253,表明具有极强的统计显著性。
- 内生-外生映射成功识别出具有高病毒传播潜力的视频(即高敏感性与高内生响应),从而支持针对性推广策略的制定。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。