Skip to main content
QUICK REVIEW

[Paper Review] Large Arabic Twitter Dataset on COVID-19

Sarah Alqurashi, Ahmad Alhindi|arXiv (Cornell University)|Apr 9, 2020
Misinformation and Its Impacts3 references33 citations
TL;DR

This paper presents the first Arabic-language Twitter dataset on COVID-19, collected since January 1, 2020, with over 3.9 million Arabic tweets and accompanying metadata, and provides data collection methods and initial statistics.

ABSTRACT

The 2019 coronavirus disease (COVID-19), emerged late December 2019 in China, is now rapidly spreading across the globe. At the time of writing this paper, the number of global confirmed cases has passed two millions and half with over 180,000 fatalities. Many countries have enforced strict social distancing policies to contain the spread of the virus. This have changed the daily life of tens of millions of people, and urged people to turn their discussions online, e.g., via online social media sites like Twitter. In this work, we describe the first Arabic tweets dataset on COVID-19 that we have been collecting since January 1st, 2020. The dataset would help researchers and policy makers in studying different societal issues related to the pandemic. Many other tasks related to behavioral change, information sharing, misinformation and rumors spreading can also be analyzed.

Motivation & Objective

  • Motivate study of Arabic-speaking public discourse and behavioral change during COVID-19.
  • Provide a large-scale Arabic Twitter dataset to analyze information spreading, sentiment, and misinformation.
  • Describe data collection methods, preprocessing plans, and initial usage guidance for researchers and policymakers.

Proposed method

  • Collect Arabic COVID-19 related tweets from Jan 1, 2020 to Apr 15, 2020 using Twitter streaming API and Tweepy.
  • Maintain full tweet objects including tweet id, username, hashtags, and geolocation when available.
  • Create a keyword list with English translations to track Arabic COVID-19 related terms and track associated tweets via streaming API.
  • Track a set of hashtags with counts and dates to monitor COVID-19 discourse in Arabic on Twitter.
  • Filter out irrelevant tweets to keep pandemic-relevant content.
  • Provide dataset access via GitHub and distribute only tweet IDs to comply with Twitter policies, noting that full objects can be retrieved with tools like Hydrator.

Experimental results

Research questions

  • RQ1What is the volume and temporal pattern of Arabic-language Twitter activity related to COVID-19 from Jan 1 to Apr 15, 2020?
  • RQ2What proportion of collected Arabic COVID-19 tweets are geotagged and how many are original tweets vs. retweets?
  • RQ3What keywords and hashtags dominate Arabic COVID-19 discourse on Twitter, and how do they evolve over time?
  • RQ4How can researchers utilize the dataset to study information sharing, misinformation, and behavioral responses in Arabic-speaking populations?
  • RQ5What are the limitations and data access considerations when using the dataset for research?

Key findings

  • Collected more than 3,934,610 Arabic tweets related to COVID-19 since Jan 1, 2020.
  • Geolocation data is available for 219 tweets, while 3,934,235 are original tweets and 375 are retweets.
  • The average daily tweet collection rate is 77,471.
  • A list of keywords and hashtags was used to track relevant tweets, with Table 1 showing keyword tracing dates and Table 2 showing hashtags and counts.
  • The dataset is hosted on GitHub and distributes only tweet IDs to comply with Twitter’s content redistribution policy; full objects can be retrieved via tools like Hydrator.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.