[논문 리뷰] This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API
이 연구는 역사적 트윗 샘플링을 위한 티커트 애카데믹 API v2의 데이터 완전성과 시간적 신뢰성을 평가하며, 전체 아카이브 액세스를 통해 거의 완전한 샘플링이 가능하다는 것을 입증한다. v2는 v1.1보다 우수한 성능을 보이며 시간이 지남에 따라 데이터 손실를 최소화하고, 데이터 최소화 원칙을 더 잘 준수하여 투명하고 재현 가능한 방법으로 더 신뢰할 수 있는 소셜미디어 연구를 지원한다.
Because of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a series of publications have studied and criticized Twitter's APIs and Twitter has partially adapted its existing data streams. The newest Twitter API for Academic Research allows to "access Twitter's real-time and historical public data with additional features and functionality that support collecting more precise, complete, and unbiased datasets." The main new feature of this API is the possibility of accessing the full archive of all historic Tweets. In this article, we will take a closer look at the Academic API and will try to answer two questions. First, are the datasets collected with the Academic API complete? Secondly, since Twitter's Academic API delivers historic Tweets as represented on Twitter at the time of data collection, we need to understand how much data is lost over time due to Tweet and account removal from the platform. Our work shows evidence that Twitter's Academic API can indeed create (almost) complete samples of Twitter data based on a wide variety of search terms. We also provide evidence that Twitter's data endpoint v2 delivers better samples than the previously used endpoint v1.1. Furthermore, collecting Tweets with the Academic API at the time of studying a phenomenon rather than creating local archives of stored Tweets, allows for a straightforward way of following Twitter's developer agreement. Finally, we will also discuss technical artifacts and implications of the Academic API. We hope that our work can add another layer of understanding of Twitter data collections leading to more reliable studies of human behavior via social media data.
연구 동기 및 목표
- 역사적 데이터 수집을 위한 티커트 애카데믹 API의 완전성과 시간적 신뢰성을 평가하기 위해.
- 과거 트윗을 검색하는 데 있어 API v2와 v1.1의 성능을 비교하기 위해.
- 삭제되거나 보호 조치가 취해진 트윗으로 인해 시간이 지남에 따라 발생하는 트윗 손실 비율을 측정하기 위해.
- 티커트의 개발자 약관과 데이터 최소화 원칙 준수 여부를 평가하기 위해.
- 연구자들이 신뢰할 수 있고 윤리적인 데이터 수집을 위한 실질적인 권고 사항을 제공하기 위해.
제안 방법
- 다양한 키워드를 사용해 여러 시간대에 걸쳐 전체 아카이브 검색을 수행하여 커버리지 평가를 수행하였다.
- 애카데믹 API를 사용해 과거 트윗을 재수집하고 원본 데이터와 비교하여 누락된 트윗을 탐지하였다.
- API에서 반환된 오류 메시지를 활용해 검색 시점에 삭제되거나 보호된 트윗을 식별하였다.
- 애카데믹 API v2와 v1.1 엔드포인트를 비교하는 제어 실험을 수행하였다.
- 기준 정확도와 완전성을 확보하기 위해 고비용의 티커트 프리미엄 API를 통해 데이터를 수집하였다.
- 동일한 쿼리를 여러 시점에 재수집하여 장기적인 트윗 가용성 감소를 분석하였다.
실험 결과
연구 질문
- RQ1티커트 애카데믹 API v2는 다양한 검색어에 대해 역사적 트윗의 거의 완전한 샘플을 제공하는가?
- RQ2API v2의 데이터 완전성은 이전의 v1.1 엔드포인트와 비교해 어떻게 다른가?
- RQ3전체 아카이브 검색 중 삭제되거나 보호된 트윗으로 인해 시간이 지남에 따라 손실되는 트윗의 비율은 얼마인가?
- RQ4연구자가 애카데믹 API를 통해 데이터를 수집할 때 티커트의 개발자 약관을 준수하기 위해 어떻게 해야 하는가?
- RQ5애카데믹 API를 통해 수집된 티커트 데이터의 신뢰성에 영향을 주는 기술적 아티팩트와 데이터 수집 아티팩트는 무엇인가?
주요 결과
- 애카데믹 API v2는 다양한 키워드와 시간대에 걸쳐 거의 완전한 트윗 데이터 샘플링을 가능하게 하며, 누락된 트윗이 극히 소수에 그친다.
- API v2는 오래된 트윗이나 자주 트윗되지 않는 콘텐츠를 검색하는 데 있어 v1.1보다 일관되게 뛰어난 성능을 보이며 데이터 완전성이 뛰어나다.
- 삭제되거나 보호된 트윗으로 인한 손실 비율이 매우 낮아 전체적으로 5% 미만으로, 시간적 신뢰성이 높음을 시사한다.
- 애카데믹 API는 명시적 필드 선택 기능을 통해 데이터 최소화 원칙 준수를 지원하여 GDPR 하에서 위험을 줄인다.
- 로컬 아카이브를 저장하는 대신 분석 시점에 데이터를 수집하는 방식은 티커트의 개발자 약관 준수를 단순화한다.
- 지연되거나 일관되지 않은 데이터 전달과 같은 기술적 아티팩트가 관찰되어 실시간 시스템과 아카이브 시스템 간의 아키텍처적 분리가 존재함을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.