[论文解读] Broker Bots: Analyzing automated activity during High Impact Events on Twitter
本文提出了一种方法,通过人工标注与机器学习相结合的方式,在重大事件期间识别和分析Twitter机器人,发现这些机器人主要聚合和传播可信来源的信息,而非传播谣言。研究发现,2013年机器人越来越多地使用IFTTT和dlvr.it等网络自动化工具,其行为总体上非恶意,基于用户特征的实时机器人分类器准确率达到85.10%。
Twitter is now an established and a widely popular news medium. Be it normal banter or a discussion on high impact events like Boston marathon blasts, February 2014 US Icestorm, etc., people use Twitter to get updates. Twitter bots have today become very common and acceptable. People are using them to get updates about emergencies like natural disasters, terrorist strikes, etc. Twitter bots provide these users a means to perform certain tasks on Twitter that are both simple and structurally repetitive. During high impact events these Twitter bots tend to provide time critical and comprehensive information. We present how bots participate in discussions and augment them during high impact events. We identify bots in high impact events for 2013: Boston blasts, February 2014 US Icestorm, Washington Navy Yard Shooting, Oklahoma tornado, and Cyclone Phailin. We identify bots among top tweeters by getting all such accounts manually annotated. We then study their activity and present many important insights. We determine the impact bots have on information diffusion during these events and how they tend to aggregate and broker information from various sources to different users. We also analyzed their tweets, list down important differentiating features between bots and non bots (normal or human accounts) during high impact events. We also show how bots are slowly moving away from traditional API based posts towards web automation platforms like IFTTT, dlvr.it, etc. Using standard machine learning, we proposed a methodology to identify bots/non bots in real time during high impact events. This study also looks into how the bot scenario has changed by comparing data from high impact events from 2013 with data from similar type of events from 2011. Lastly, we also go through an in-depth analysis of Twitter bots who were active during 2013 Boston Marathon Blast.
研究动机与目标
- 理解自动化账户(机器人)在Twitter重大事件中信息传播中的作用。
- 通过人工标注与机器学习,在危机事件中识别并区分机器人与人类账户。
- 分析2011年至2013年间机器人行为及发布机制的变化,特别是从基于API的发布方式向网络自动化平台的转变。
- 评估机器人在突发新闻事件中对谣言传播与信息扩散的影响。
- 利用用户特征与时间特征,构建基于用户层面的实时机器人检测模型,以支持危机监控。
提出的方法
- 通过Twitter API收集2013年五起重大事件的数据,包括波士顿马拉松爆炸、Icestorm事件、华盛顿海军工厂枪击案、俄克拉荷马龙卷风和气旋Phailin。
- 通过人工标注方式,对这些事件中约1,000名顶级发帖者进行机器人或非机器人分类,共识别出377个机器人和115个非机器人。
- 对标注数据进行丰富,包括关注者/粉丝网络、个人资料描述、推文来源、URL链接及发布时间。
- 使用WEKA基于用户特征(如关注者/粉丝比例、个人资料关键词等)训练机器学习分类器,以区分机器人与人类,准确率达到85.10%。
- 将2013年机器人行为与2011年危机事件数据进行对比,识别发布机制与信息来源的变化。
- 对个别机器人进行深入分析,包括互动模式与内容模式,以理解其在信息中介中的作用。
实验结果
研究问题
- RQ1机器人在Twitter重大事件期间如何参与信息传播?
- RQ2危机事件中,机器人与非机器人账户的显著特征是什么?
- RQ32011年至2013年间,机器人发布方式发生了怎样的演变,特别是平台使用的变化?
- RQ4机器人在突发新闻事件中对谣言传播的贡献程度如何?
- RQ5能否在事件进行期间,基于用户层面特征有效构建实时机器人检测模型?
主要发现
- 2013年的机器人主要使用IFTTT和dlvr.it等网络自动化平台,标志着与2011年传统基于API的发布方式相比的显著转变。
- 基于用户特征的机器学习分类器在区分机器人与非机器人方面达到85.10%的准确率,而时间特征表现较差。
- 机器人相较于非机器人具有显著更高的关注者/粉丝比例,中位数比率为1.53(机器人)对比0.31(非机器人),表明其具有广播式行为特征。
- 机器人极少传播谣言;即使传播,也存在显著延迟,表明其依赖经验证或可信的信息源。
- 机器人个人资料中频繁出现“news”(新闻)、“updates”(更新)、“breaking”(突发)、“source”(来源)等关键词,其关注者也常包含类似关键词,表明存在网络化信息中介模式。
- 在波士顿马拉松爆炸事件中,机器人及其关注者形成的网络呈现中心-辐射结构,机器人作为中介,将来自认证来源的信息分发给广大受众。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。