Skip to main content
QUICK REVIEW

[论文解读] It Doesn't Break Just on Twitter. Characterizing Facebook content During Real World Events

Prateek Dewan, Ponnurangam Kumaraguru|arXiv (Cornell University)|May 19, 2014
Misinformation and Its Impacts参考文献 23被引用 4
一句话总结

本研究对比了16起现实事件中Facebook与Twitter的内容,发现30.56%的公开Facebook内容也出现在Twitter上,且Facebook平均在事件发生后12分钟内发布新闻。利用文体特征,作者在Facebook上实现了超过99%的垃圾信息分类准确率,在Twitter上达到98%,揭示了尽管Facebook默认设置为私密,但其内容仍存在显著重叠与实时传播潜力。

ABSTRACT

Multiple studies in the past have analyzed the role and dynamics of the Twitter social network during real world events. However, little work has explored the content of other social media services, or compared content across two networks during real world events. We believe that social media platforms like Facebook also play a vital role in disseminating information on the Internet during real world events. In this work, we study and characterize the content posted on the world's biggest social network, Facebook, and present a comparative analysis of Facebook and Twitter content posted during 16 real world events. Contrary to existing notion that Facebook is used mostly as a private network, our findings reveal that more than 30% of public content that was present on Facebook during these events, was also present on Twitter. We then performed qualitative analysis on the content spread by the most active users during these events, and found that over 10% of the most active users on both networks post spam content. We used stylometric features from Facebook posts and tweet text to classify this spam content, and were able to achieve an accuracy of over 99% for Facebook, and over 98% for Twitter. This work is aimed at providing researchers with an overview of Facebook content during real world events, and serve as basis for more in-depth exploration of its various aspects like information quality, and credibility during real world events.

研究动机与目标

  • 调查在现实事件中,Facebook是否作为公共信息的重要来源,挑战其主要为私人网络的假设。
  • 比较Facebook与Twitter在重大全球事件中信息传播的内容、时效性与特征差异。
  • 分析两个平台上的垃圾信息行为,并评估仅使用文本文体特征检测垃圾信息的可行性。
  • 评估Facebook在危机期间作为可信且及时信息来源的潜力,与Twitter相媲美。

提出的方法

  • 使用与事件相关的数据采集框架,收集16起现实事件期间公开可获取的Facebook与Twitter内容。
  • 提取并分析两个平台的文本内容、话题标签和URL,以比较内容重叠与时间动态。
  • 从帖子中提取文体特征(如词频、标点符号使用、字符级模式)以训练垃圾信息检测分类器。
  • 应用监督式机器学习模型,仅基于文本特征对两个网络的垃圾信息进行分类,实现高准确率。
  • 分析内容与垃圾信息帖子的时间模式,以检测爆发性行为与自动化行为。
  • 对活跃度最高的用户进行定性分析,以理解内容传播与垃圾信息传播机制。

实验结果

研究问题

  • RQ1在现实事件期间,Facebook上的公开内容与Twitter内容的重叠程度如何?
  • RQ2与Twitter相比,Facebook在传播现实事件信息方面速度如何?
  • RQ3在现实事件期间,Facebook上的垃圾信息普遍程度与性质如何?与Twitter相比有何差异?
  • RQ4仅使用文体特征能否准确检测Facebook与Twitter上的垃圾信息?
  • RQ5在重大事件期间,Facebook与Twitter上活跃度最高用户的行为模式有何不同?

主要发现

  • 在16起现实事件期间,30.56%的公开Facebook内容也出现在Twitter上,表明两个平台间存在显著的内容重叠。
  • Facebook在事件发生后平均11.2分钟内即发布新闻,显著快于此前报道的Twitter时间线。
  • 在研究的事件中,两个平台最活跃用户中超过10%发布了垃圾信息。
  • 仅使用文体特征进行垃圾信息检测,在Facebook上准确率达99.1%,在Twitter上达98.3%,表明基于文本的分类方法具有极高有效性。
  • Facebook上的垃圾信息表现出爆发式发布行为,而Twitter上的垃圾信息则呈现更稳定、自动化的发布模式,表明两个平台的垃圾信息策略存在差异。
  • 尽管Facebook默认隐私设置为私密,但其公开内容总量(每日13.3亿条)远超Twitter(每日5亿条),使其在事件期间成为丰富信息来源。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。