Skip to main content
QUICK REVIEW

[论文解读] This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API

Jürgen Pfeffer, Angelina Mooseder|arXiv (Cornell University)|Apr 4, 2022
Web Data Mining and Analysis被引用 14
一句话总结

本研究评估了Twitter学术版API v2在数据完整性和时间可靠性方面的表现,证明其通过全档案访问功能可实现对历史推文的近乎完整采样。研究发现v2优于v1.1,随时间推移的数据丢失更少,并且更符合数据最小化原则,支持采用透明、可复现方法进行更可靠的社交媒体研究。

ABSTRACT

Because of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a series of publications have studied and criticized Twitter's APIs and Twitter has partially adapted its existing data streams. The newest Twitter API for Academic Research allows to "access Twitter's real-time and historical public data with additional features and functionality that support collecting more precise, complete, and unbiased datasets." The main new feature of this API is the possibility of accessing the full archive of all historic Tweets. In this article, we will take a closer look at the Academic API and will try to answer two questions. First, are the datasets collected with the Academic API complete? Secondly, since Twitter's Academic API delivers historic Tweets as represented on Twitter at the time of data collection, we need to understand how much data is lost over time due to Tweet and account removal from the platform. Our work shows evidence that Twitter's Academic API can indeed create (almost) complete samples of Twitter data based on a wide variety of search terms. We also provide evidence that Twitter's data endpoint v2 delivers better samples than the previously used endpoint v1.1. Furthermore, collecting Tweets with the Academic API at the time of studying a phenomenon rather than creating local archives of stored Tweets, allows for a straightforward way of following Twitter's developer agreement. Finally, we will also discuss technical artifacts and implications of the Academic API. We hope that our work can add another layer of understanding of Twitter data collections leading to more reliable studies of human behavior via social media data.

研究动机与目标

  • 评估Twitter学术版API在历史数据采集中的完整性和时间可靠性。
  • 比较API v2与v1.1在检索历史推文方面的性能表现。
  • 测量由于删除或设为保护状态而导致的推文随时间推移的丢失率。
  • 评估其与Twitter开发者协议及数据最小化原则的合规性。
  • 为研究人员提供可操作的建议,以实现可靠且符合伦理的数据采集。

提出的方法

  • 通过在多个时间段使用多样化关键词执行全档案搜索,以评估数据覆盖范围。
  • 使用学术版API重新收集历史推文,并与原始数据对比以检测缺失的推文。
  • 利用API返回的错误信息识别在检索时已被删除或设为保护状态的推文。
  • 通过受控实验比较通过学术版API v2和v1.1端点收集数据的表现。
  • 通过成本较高的Twitter Premium API收集数据,作为准确性和完整性的基准。
  • 通过在多个时间点重新收集相同查询,分析推文可用性随时间推移的衰减情况。

实验结果

研究问题

  • RQ1Twitter学术版API v2在多样化搜索关键词下,对历史推文的采样完整度如何?
  • RQ2与之前的v1.1端点相比,API v2的数据完整性表现如何?
  • RQ3在全档案搜索过程中,由于删除或设为保护状态,有多少比例的推文随时间推移而丢失?
  • RQ4研究人员如何确保通过学术版API收集数据时符合Twitter的开发者协议?
  • RQ5哪些技术特征和数据采集过程中的特征会影响通过学术版API收集的Twitter数据的可靠性?

主要发现

  • 学术版API v2在多样化关键词和时间段下可实现近乎完整的Twitter数据采样,缺失推文极少。
  • API v2在数据完整性方面始终优于v1.1,尤其在检索较旧或较少被提及的内容时表现更优。
  • 仅有极小比例的推文(不足5%)因删除或设为保护状态而随时间推移丢失,表明其具有较强的时间可靠性。
  • 学术版API通过支持显式字段选择,有助于符合数据最小化原则,降低在GDPR下的风险。
  • 在分析时即时收集数据而非存储本地档案,可简化对Twitter开发者协议的遵守。
  • 观察到延迟或不一致的数据交付等技术特征,表明实时系统与归档系统之间存在架构上的分离。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。