[论文解读] Agents, Bookmarks and Clicks: A topical model of Web traffic
本文提出ABC代理模型,通过整合书签、返回键导航与主题兴趣,解释个体浏览行为与整体网络流量模式。该模型成功再现了页面/链接流量、会话大小、熵值及起始页面受欢迎程度的经验分布,通过非马尔可夫机制捕捉现实网络导航的异质性,优于PageRank与BookRank。
Analysis of aggregate and individual Web traffic has shown that PageRank is a poor model of how people navigate the Web. Using the empirical traffic patterns generated by a thousand users, we characterize several properties of Web traffic that cannot be reproduced by Markovian models. We examine both aggregate statistics capturing collective behavior, such as page and link traffic, and individual statistics, such as entropy and session size. No model currently explains all of these empirical observations simultaneously. We show that all of these traffic patterns can be explained by an agent-based model that takes into account several realistic browsing behaviors. First, agents maintain individual lists of bookmarks (a non-Markovian memory mechanism) that are used as teleportation targets. Second, agents can retreat along visited links, a branching mechanism that also allows us to reproduce behaviors such as the use of a back button and tabbed browsing. Finally, agents are sustained by visiting novel pages of topical interest, with adjacent pages being more topically related to each other than distant ones. This modulates the probability that an agent continues to browse or starts a new session, allowing us to recreate heterogeneous session lengths. The resulting model is capable of reproducing the collective and individual behaviors we observe in the empirical data, reconciling the narrowly focused browsing patterns of individual users with the extreme heterogeneity of aggregate traffic measurements. This result allows us to identify a few salient features that are necessary and sufficient to interpret the browsing patterns observed in our data. In addition to the descriptive and explanatory power of such a model, our results may lead the way to more sophisticated, realistic, and effective ranking and crawling algorithms.
研究动机与目标
- 解决PageRank等马尔可夫模型在解释现实网络浏览行为方面的局限性。
- 调和个体会话的狭隘焦点与整体流量测量中观察到的极端异质性之间的矛盾。
- 开发一个考虑个体记忆(书签)、导航机制(返回键)与主题兴趣的现实代理模型。
- 同时解释个体层面指标(熵值、会话大小)与整体层面流量模式(页面/链接受欢迎程度)。
- 通过建模人类导航的关键特征,为改进排名与爬取算法提供基础。
提出的方法
- 代理维护个性化的书签列表作为跳转目标,实现会话边界与真实的起始页面受欢迎程度。
- 通过返回键机制,代理可沿已访问链接回退,模拟分页浏览与分支导航模式。
- 通过链接内容相似的页面,建模主题性,提高在相关主题上继续浏览的概率。
- 浏览决策受主题兴趣调节:若内容相关则继续浏览,否则开启新会话。
- 模型采用非均匀、主题感知的链接选择过程,反映用户探索相关内容的实际方式。
- 使用印第安纳大学1,000名用户的实证数据,校准并验证模型与真实流量模式的一致性。
实验结果
研究问题
- RQ1为何PageRank与BookRank无法预测个体用户流量熵值与会话大小分布?
- RQ2哪些非马尔可夫机制是解释聚焦个体浏览与异质性整体流量共存所必需的?
- RQ3书签、返回键使用与主题兴趣如何协同再现个体与集体层面的观测流量模式?
- RQ4具备记忆、分支与主题性的代理模型在多大程度上优于无记忆模型如PageRank?
- RQ5主题局部性在塑造会话长度与网络导航中的用户参与度方面发挥何种作用?
主要发现
- ABC模型在再现页面流量、链接流量与会话起始页面受欢迎程度的经验分布方面,优于PageRank与BookRank。
- 个体用户熵值(衡量浏览多样性)由ABC模型捕捉得比PageRank或BookRank更佳,表明其对聚焦但多样的用户行为建模更优。
- 该模型成功再现了真实用户数据中观察到的广泛且异质的会话大小与深度分布,而简单模型无法捕捉此特征。
- 仅靠书签可修正PageRank中观察到的整体流量不匹配问题,但不足以建模个体会话的异质性。
- 返回键机制对再现分叉导航模式与分页浏览行为至关重要,与实证数据一致。
- 链接选择中的主题局部性解释了用户为何持续在相关主题上浏览,直接影响会话长度与深度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。