Skip to main content
QUICK REVIEW

[论文解读] VIP: Incorporating Human Cognitive Biases in a Probabilistic Model of Retweeting

Jeon-Hyung Kang, Kristina Lermam|arXiv (Cornell University)|Feb 2, 2015
Complex Network Analysis Techniques参考文献 14被引用 9
一句话总结

本文提出VIP模型,一种概率模型,通过整合人类认知偏见——信息在信息流中的可见性(因位置而异)、项目适应度(病毒式传播潜力)以及与用户兴趣的相关性——来提升Twitter上转发行为的预测能力。通过基于用户信息负荷估算可见性,该模型显著优于忽略可见性的基线模型,证明认知因素能增强信息传播的可预测性。

ABSTRACT

Information spread in social media depends on a number of factors, including how the site displays information, how users navigate it to find items of interest, users' tastes, and the `virality' of information, i.e., its propensity to be adopted, or retweeted, upon exposure. Probabilistic models can learn users' tastes from the history of their item adoptions and recommend new items to users. However, current models ignore cognitive biases that are known to affect behavior. Specifically, people pay more attention to items at the top of a list than those in lower positions. As a consequence, items near the top of a user's social media stream have higher visibility, and are more likely to be seen and adopted, than those appearing below. Another bias is due to the item's fitness: some items have a high propensity to spread upon exposure regardless of the interests of adopting users. We propose a probabilistic model that incorporates human cognitive biases and personal relevance in the generative model of information spread. We use the model to predict how messages containing URLs spread on Twitter. Our work shows that models of user behavior that account for cognitive factors can better describe and predict user behavior in social media.

研究动机与目标

  • 解决现有社交推荐模型忽略影响信息传播中用户行为的认知偏见这一缺陷。
  • 通过将可见性建模为用户信息负荷的函数以及信息在信息流中的位置,提升转发行为预测的准确性。
  • 解耦项目适应度、个人相关性和可见性在Twitter信息级联传播中的贡献。
  • 证明整合认知偏见可提升社交媒体中用户采纳行为模型的准确性和鲁棒性。

提出的方法

  • VIP模型采用生成式概率框架,联合建模项目可见性、适应度以及与用户兴趣的相关性。
  • 可见性通过用户信息负荷间接估算——即用户为到达某一项目需检查的新消息数量,该值由关注者数量和访问频率推导得出。
  • 个人相关性通过隐含主题模型计算,基于采纳历史推断用户和项目的话题分布。
  • 项目适应度被建模为独立于用户兴趣的潜在采纳倾向,以捕捉病毒式传播效应。
  • 模型通过在Twitter URL转发数据上使用最大似然估计,联合估计这三个组成部分。
  • 通过对比仅使用相关性、仅使用适应度或随机选择的基线模型,评估VIP的预测性能。

实验结果

研究问题

  • RQ1在考虑个人相关性和适应度之后,由信息在用户信息流中位置决定的可见性,能在多大程度上解释转发行为的方差?
  • RQ2项目适应度、个人相关性和可见性如何共同影响Twitter上信息级联的规模?
  • RQ3是否能够设计出一种整合认知偏见(如可见性和适应度)的模型,使其优于仅依赖用户兴趣画像的传统社交推荐模型?
  • RQ4可见性、适应度和相关性对不同类型内容(如名人新闻与小众爱好内容)的相对贡献如何变化?

主要发现

  • 通过用户信息负荷估算的可见性,显著提升了忽略位置偏见的模型的预测准确性。
  • VIP模型在所有用户活跃度水平下均优于基线模型,尤其在低活跃度用户中提升最大,因此时个人相关性不确定性更高。
  • 对于高质量项目(高适应度与高相关性),可见性可解释级联规模的额外方差:越可见的项目传播范围越广。
  • 40%的URL未显示适应度与级联规模之间的相关性,表明仅靠适应度不足以预测病毒式传播。
  • ‘Jay-Z音乐视频’的适应度是‘Ian Somerhalder基金会’的六倍,但其预期可见性仅为一半,却获得了相似的转发数,表明适应度可弥补较低的可见性。
  • 项目适应度与级联规模之间存在0.85的统计显著相关性,但仅适用于部分URL,凸显多因素模型的必要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。