[Paper Review] Extracting COVID-19 Events from Twitter
This paper introduces a manually annotated corpus of 7,500 Twitter tweets labeled with COVID-19 events such as positive test results and denied testing access. Using this corpus, the authors demonstrate effective automatic extraction of event mentions through slot-filling, enabling structured representation of self-reported pandemic-related experiences, with tools and data made publicly available.
We present a corpus of 7,500 tweets annotated with COVID-19 events, including positive test results, denied access to testing, and more. We show that our corpus enables automatic identification of COVID-19 events mentioned in Twitter with text spans that fill a set of pre-defined slots for each event. We also present analyses on the self-reporting cases and user's demographic information. We will make our annotated corpus and extraction tools available for the research community to use upon publication at this https URL
Motivation & Objective
- To create a high-quality, manually annotated corpus of Twitter posts reporting COVID-19-related events.
- To enable automatic identification of specific event types—such as positive test results and denied testing access—using text spans in predefined slots.
- To analyze self-reported cases and user demographic patterns in social media during the pandemic.
- To support future NLP research by releasing the annotated corpus and extraction tools to the public.
Proposed method
- Manual annotation of 7,500 tweets with predefined event types and slot-filling structures for each event.
- Definition of a set of pre-defined event slots (e.g., event type, date, location) to standardize event representation.
- Application of sequence labeling and event extraction techniques to identify spans in text corresponding to each slot.
- Use of statistical and qualitative analysis to examine patterns in self-reported cases and user demographics.
- Development of tools for event extraction based on the annotated corpus for downstream use.
Experimental results
Research questions
- RQ1How accurately can event types such as positive test results and denied access to testing be identified in social media text?
- RQ2To what extent do self-reported cases in Twitter align with official case data or demographic trends?
- RQ3What patterns emerge in user demographics among individuals reporting COVID-19 events on Twitter?
- RQ4Can a structured slot-filling approach effectively extract meaningful event information from informal social media language?
Key findings
- The annotated corpus enables accurate automatic identification of COVID-19 events through slot-filling, supporting structured event extraction.
- A significant number of self-reported cases were found in the corpus, reflecting real-time public health reporting via social media.
- Demographic patterns in user reports suggest variations in reporting behavior across different user groups.
- The corpus and associated tools are publicly released to support ongoing research in NLP and public health informatics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.