[论文解读] Large Arabic Twitter Dataset on COVID-19
本文介绍了首个阿拉伯语 COVID-19 推特数据集,自 January 1, 2020 起收集,包含超过 3.9 million 阿拉伯语推文及相关元数据,并提供数据收集方法与初步统计信息。
The 2019 coronavirus disease (COVID-19), emerged late December 2019 in China, is now rapidly spreading across the globe. At the time of writing this paper, the number of global confirmed cases has passed two millions and half with over 180,000 fatalities. Many countries have enforced strict social distancing policies to contain the spread of the virus. This have changed the daily life of tens of millions of people, and urged people to turn their discussions online, e.g., via online social media sites like Twitter. In this work, we describe the first Arabic tweets dataset on COVID-19 that we have been collecting since January 1st, 2020. The dataset would help researchers and policy makers in studying different societal issues related to the pandemic. Many other tasks related to behavioral change, information sharing, misinformation and rumors spreading can also be analyzed.
研究动机与目标
- Motivate study of Arabic-speaking public discourse and behavioral change during COVID-19.
- Provide a large-scale Arabic Twitter dataset to analyze information spreading, sentiment, and misinformation.
- Describe data collection methods, preprocessing plans, and initial usage guidance for researchers and policymakers.
提出的方法
- Collect Arabic COVID-19 related tweets from Jan 1, 2020 to Apr 15, 2020 using Twitter streaming API and Tweepy.
- Maintain full tweet objects including tweet id, username, hashtags, and geolocation when available.
- Create a keyword list with English translations to track Arabic COVID-19 related terms and track associated tweets via streaming API.
- Track a set of hashtags with counts and dates to monitor COVID-19 discourse in Arabic on Twitter.
- Filter out irrelevant tweets to keep pandemic-relevant content.
- Provide dataset access via GitHub and distribute only tweet IDs to comply with Twitter policies, noting that full objects can be retrieved with tools like Hydrator.
实验结果
研究问题
- RQ1What is the volume and temporal pattern of Arabic-language Twitter activity related to COVID-19 from Jan 1 to Apr 15, 2020?
- RQ2What proportion of collected Arabic COVID-19 tweets are geotagged and how many are original tweets vs. retweets?
- RQ3What keywords and hashtags dominate Arabic COVID-19 discourse on Twitter, and how do they evolve over time?
- RQ4How can researchers utilize the dataset to study information sharing, misinformation, and behavioral responses in Arabic-speaking populations?
- RQ5What are the limitations and data access considerations when using the dataset for research?
主要发现
- Collected more than 3,934,610 Arabic tweets related to COVID-19 since Jan 1, 2020.
- Geolocation data is available for 219 tweets, while 3,934,235 are original tweets and 375 are retweets.
- The average daily tweet collection rate is 77,471.
- A list of keywords and hashtags was used to track relevant tweets, with Table 1 showing keyword tracing dates and Table 2 showing hashtags and counts.
- The dataset is hosted on GitHub and distributes only tweet IDs to comply with Twitter’s content redistribution policy; full objects can be retrieved via tools like Hydrator.
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。