[論文レビュー] This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API
本研究では、Twitter Academic API v2 のデータ完全性および時間的信頼性を評価し、フルアーカイブアクセスを活用することで、歴史的ツイートのほぼ完全なサンプリングが可能であることを示している。v2 は v1.1 を上回り、時間経過に伴うデータ損失を最小限に抑え、データ最小化原則への適合性も向上しており、透明性と再現可能性の高い方法で、より信頼性の高いソーシャルメディア研究を支援する。
Because of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a series of publications have studied and criticized Twitter's APIs and Twitter has partially adapted its existing data streams. The newest Twitter API for Academic Research allows to "access Twitter's real-time and historical public data with additional features and functionality that support collecting more precise, complete, and unbiased datasets." The main new feature of this API is the possibility of accessing the full archive of all historic Tweets. In this article, we will take a closer look at the Academic API and will try to answer two questions. First, are the datasets collected with the Academic API complete? Secondly, since Twitter's Academic API delivers historic Tweets as represented on Twitter at the time of data collection, we need to understand how much data is lost over time due to Tweet and account removal from the platform. Our work shows evidence that Twitter's Academic API can indeed create (almost) complete samples of Twitter data based on a wide variety of search terms. We also provide evidence that Twitter's data endpoint v2 delivers better samples than the previously used endpoint v1.1. Furthermore, collecting Tweets with the Academic API at the time of studying a phenomenon rather than creating local archives of stored Tweets, allows for a straightforward way of following Twitter's developer agreement. Finally, we will also discuss technical artifacts and implications of the Academic API. We hope that our work can add another layer of understanding of Twitter data collections leading to more reliable studies of human behavior via social media data.
研究の動機と目的
- 歴史的データ収集における Twitter Academic API の完全性および時間的信頼性を評価すること。
- API v2 と v1.1 の間で、歴史的ツイートの取得性能を比較すること。
- 削除または保護されたツイートによる、時間経過に伴うツイート損失率を測定すること。
- Twitter の開発者契約およびデータ最小化原則への適合性を評価すること。
- 研究者が信頼性があり倫理的なデータ収集が行えるよう、実行可能な推奨事項を提供すること。
提案手法
- 多様なキーワードを用いたフルアーカイブ検索を実施し、カバレッジを評価した。
- Academic API を用いて歴史的ツイートを再収集し、元のデータと照合することで欠落ツイートを特定した。
- API からのエラーメッセージを活用し、取得時刻に削除または保護されたツイートを同定した。
- Academic API v2 と v1.1 のエンドポイントを比較する制御実験を実施した。
- ベンチマークとして、高コストな Twitter Premium API を用いてデータを収集した。
- 複数の時系列で同じクエリを再収集することで、ツイートの可用性の長期的低下を分析した。
実験結果
リサーチクエスチョン
- RQ1Twitter Academic API v2 は、多様な検索キーワードに対して、歴史的ツイートのほぼ完全なサンプルをどの程度提供するか?
- RQ2API v2 のデータ完全性は、以前の v1.1 エンドポイントと比べてどのように異なるか?
- RQ3フルアーカイブ検索において、削除または保護のため時間経過に伴い失われるツイートの割合はどの程度か?
- RQ4研究者が Academic API を通じてデータを収集する際、どのようにして Twitter の開発者契約に準拠できるか?
- RQ5Academic API を通じて収集される Twitter データの信頼性に影響を与える技術的アーティファクトおよびデータ収集アーティファクトは何か?
主な発見
- Academic API v2 は、多様なキーワードおよび時間帯において、ほぼ完全な Twitter データサンプリングを可能にし、欠落ツイートは最小限に抑えられる。
- API v2 は、とくに古いツイートや頻度の低いツイートを収集する際、v1.1 よりも一貫して高いデータ完全性を示す。
- 削除または保護のための時間経過によるツイート損失は、わずか 5% 未満にとどまり、強い時間的信頼性を示している。
- 明示的なフィールド選択が可能であるため、Academic API はデータ最小化原則への適合を支援し、GDPR におけるリスク低減に寄与する。
- ローカルアーカイブを保存するのではなく、分析時におけるデータ収集を行うことで、Twitter の開発者契約への適合が簡素化される。
- 遅延や一貫性の欠如といった技術的アーティファクトが観察され、リアルタイム系とアーカイブ系のシステムがアーキテクチャ的に分離されている可能性が示唆された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。