Skip to main content
QUICK REVIEW

[论文解读] Variation across Scales: Measurement Fidelity under Twitter Data Sampling

Siqi Wu, Marian-Andrei Rizoiu|arXiv (Cornell University)|Mar 21, 2020
Complex Network Analysis Techniques参考文献 34被引用 5
一句话总结

本文研究了Twitter的筛选流API采样在多个时间尺度和主题下对数据质量的影响。通过使用子爬虫重建完整的推文流,本文表明Twitter的速率限制消息能准确指示缺失的推文数量,并证明采样会扭曲实体频率、网络结构和转发级联动态——同时提供了从采样数据中估计真实统计数据的方法。

ABSTRACT

A comprehensive understanding of data quality is the cornerstone of measurement studies in social media research. This paper presents in-depth measurements on the effects of Twitter data sampling across different timescales and different subjects (entities, networks, and cascades). By constructing complete tweet streams, we show that Twitter rate limit message is an accurate indicator for the volume of missing tweets. Sampling also differs significantly across timescales. While the hourly sampling rate is influenced by the diurnal rhythm in different time zones, the millisecond level sampling is heavily affected by the implementation choices. For Twitter entities such as users, we find the Bernoulli process with a uniform rate approximates the empirical distributions well. It also allows us to estimate the true ranking with the observed sample data. For networks on Twitter, their structures are altered significantly and some components are more likely to be preserved. For retweet cascades, we observe changes in distributions of tweet inter-arrival time and user influence, which will affect models that rely on these features. This work calls attention to noises and potential biases in social data, and provides a few tools to measure Twitter sampling effects.

研究动机与目标

  • 理解Twitter的筛选流采样如何影响不同时间尺度和主题下的数据完整性和测量保真度。
  • 调查Twitter的速率限制消息是否可靠地指示了缺失推文的数量。
  • 评估采样对关键社交媒体度量的影响:实体频率、排名、网络结构和转发级联。
  • 开发从采样数据中估计真实统计数据(如频率、排名)的方法,而无需访问完整数据。
  • 为研究人员提供工具和实证证据,以减轻Twitter采样机制引入的偏差。

提出的方法

  • 使用多个子爬虫构建完整的推文流,追踪先前研究中的关键词和语言,避免使用昂贵的Firehose服务。
  • 通过将完整数据流与Twitter筛选流API的采样数据进行比较,测量采样影响。
  • 将速率限制消息用作缺失推文数量的代理,并通过实证对比验证其准确性。
  • 将实体采样建模为具有均匀速率的伯努利过程,表明其能很好地拟合经验分布。
  • 应用统计推断,从采样观测中估计真实实体频率和排名。
  • 分析转发网络和级联的结构变化,重点关注互到达时间及潜在传播范围,使用CCDF(互补累积分布函数)。

实验结果

研究问题

  • RQ1Twitter的速率限制消息在多大程度上准确反映了筛选流中缺失推文的数量?
  • RQ2采样在不同时间尺度(每小时、毫秒级)下如何变化,其时间变化的原因是什么?
  • RQ3采样在多大程度上扭曲了实体频率和排名,是否能从样本中估计真实值?
  • RQ4采样影响如何改变转发网络中的网络结构和级联动态?
  • RQ5采样对扩散模型特征(如互到达时间与用户影响力)有何影响?

主要发现

  • Twitter的速率限制消息能紧密逼近缺失推文的数量,是数据丢失的可靠指标。
  • 采样率在不同时间尺度上存在显著差异:每小时采样受昼夜节律影响,而毫秒级采样则受实现选择的影响。
  • 具有均匀速率的伯努利过程能准确建模实体采样,从而可从样本数据中估计真实实体频率和排名。
  • 完整数据中转发的中位互到达时间从22.9秒增加到采样数据中的105.7秒——超过4.6倍,扭曲了病毒式传播的评估。
  • 级联的相对潜在传播范围被低估:50%的级联在采样数据中传播范围不足真实值的54.4%,且早期预测的偏差更大。
  • 对于至少有50次转发的级联,样本中仅捕获了完整级联数量的29.6%,且每级联的平均转发数降至真实值的70.2%。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。