[Paper Review] ArCOV-19: The First Arabic COVID-19 Twitter Dataset with Propagation Networks
ArCOV-19 is the first publicly available Arabic COVID-19 Twitter dataset (Jan 2020–Jan 2021) with both source tweets and propagation networks for the most popular subset.
In this paper, we present ArCOV-19, an Arabic COVID-19 Twitter dataset that spans one year, covering the period from 27th of January 2020 till 31st of January 2021. ArCOV-19 is the first publicly-available Arabic Twitter dataset covering COVID-19 pandemic that includes about 2.7M tweets alongside the propagation networks of the most-popular subset of them (i.e., most-retweeted and -liked). The propagation networks include both retweets and conversational threads (i.e., threads of replies). ArCOV-19 is designed to enable research under several domains including natural language processing, information retrieval, and social computing. Preliminary analysis shows that ArCOV-19 captures rising discussions associated with the first reported cases of the disease as they appeared in the Arab world. In addition to the source tweets and propagation networks, we also release the search queries and language-independent crawler used to collect the tweets to encourage the curation of similar datasets.
Motivation & Objective
- Curate a year-long Arabic Twitter dataset focused on COVID-19 to enable NLP, IR, and social computing research.
- Provide propagation networks (retweets and conversational threads) for the most-popular subset of tweets.
- Release crawl methodology, queries, and tools to facilitate replication and extension by other researchers.
Proposed method
- Collect source tweets daily via a language-specific Twitter search crawler using manually crafted Arabic queries.
- Exclude retweets and duplicates; use GetOldTweets3 to recover old tweets when the search API window limits were reached.
- Identify the top 1,000 most popular tweets per day by a popularity score (sum of retweets and favorites) and collect their retweets and replies to form propagation networks.
- Release: source tweet IDs, search queries, top subset IDs, and propagation networks (retweets and conversational threads) for researchers.
- Analyze temporal, geographical, and topical characteristics of the dataset to validate coverage of the early COVID-19 period in the Arab world.
Experimental results
Research questions
- RQ1What are the temporal dynamics of Arabic COVID-19 discourse on Twitter from Jan 2020 to Jan 2021?
- RQ2What is the geographical distribution of Arabic COVID-19 tweets and how does it reflect outbreak timelines in Arab countries?
- RQ3What topics, hashtags, and domains dominate Arabic COVID-19 discussions and how do they evolve over time?
- RQ4How do propagation networks (retweets and replies) for the most popular tweets illuminate information diffusion and potential misinformation dynamics?
Key findings
- The dataset comprises about 2.7 million source tweets from over 690k users; 18.66% are from verified users and 25.40% include URLs.
- Top subset includes 370,132 tweets (13.84% of source tweets); total retweets in propagation networks reach 7,925,821 and replies reach 1,476,950 (as of April 2020 for replies).
- Geolocated tweets (place attribute) number 60,873 (2.28%), posted by 24,072 users; geotagged tweets (coordinates) number 2,078 (0.08%), from 102 countries mainly in the Arab world (92.75%).
- Saudi Arabia contributes about 41.7% of geolocated content, Kuwait about 9.6%, reflecting early outbreak timelines and country activity.
- URLs in top subset are dominated by news websites from Egypt, Saudi Arabia, and UAE; YouTube videos are among frequently shared media.
- Propagation patterns show many tweets receiving more than 100 retweets, with some exceeding 10k, indicating strong diffusion for top tweets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.