[Paper Review] Generalizable Natural Language Processing Framework for Migraine Reporting from Social Media
This paper presents a generalizable NLP framework for detecting self-reported migraine discussions in social media, using a platform-agnostic text classification system trained on 5,750 Twitter and 302 Reddit posts. It achieves F1 scores of 0.90 on Twitter and 0.93 on Reddit, enabling robust analysis of migraine-related therapies and patient sentiments from real-world online discourse.
Migraine is a highly prevalent and disabling neurological disorder. However, information about migraine management in real-world settings is limited to traditional health information sources. In this paper, we (i) verify that there is substantial migraine-related chatter available on social media (Twitter and Reddit), self-reported by those with migraine; (ii) develop a platform-independent text classification system for automatically detecting self-reported migraine-related posts, and (iii) conduct analyses of the self-reported posts to assess the utility of social media for studying this problem. We manually annotated 5750 Twitter posts and 302 Reddit posts, and used them for training and evaluating supervised machine learning methods. Our best system achieved an F<sub>1</sub> score of 0.90 on Twitter and 0.93 on Reddit. Analysis of information posted by our 'migraine cohort' revealed the presence of a plethora of relevant information about migraine therapies and sentiments associated with them. Our study forms the foundation for conducting an in-depth analysis of migraine-related information using social media data.
Motivation & Objective
- To verify the presence of substantial migraine-related discussions on social media platforms like Twitter and Reddit.
- To develop a platform-independent text classification system for identifying self-reported migraine posts.
- To analyze the content of detected migraine posts to extract insights on therapies and patient sentiments.
Proposed method
- Manual annotation of 5,750 Twitter posts and 302 Reddit posts to create a gold-standard dataset for training and evaluation.
- Training a text classification model using a transfer learning approach on diverse social media text to ensure platform independence.
- Employing a multi-class classification strategy to detect migraine-related content with high precision and recall.
- Validating model performance using F1 score as the primary evaluation metric across both Twitter and Reddit datasets.
- Applying the trained model to detect and extract information on migraine therapies and associated patient sentiments from unstructured social media text.
Experimental results
Research questions
- RQ1How much self-reported migraine-related content exists on Twitter and Reddit?
- RQ2Can a single NLP model generalize across different social media platforms to detect migraine-related posts?
- RQ3What types of migraine-related information, including therapies and sentiments, are commonly expressed in social media posts?
- RQ4How accurate is the proposed framework in identifying self-reported migraine content across platforms?
- RQ5What insights can be derived from analyzing the linguistic and affective content of detected migraine posts?
Key findings
- The framework achieved an F1 score of 0.90 on Twitter and 0.93 on Reddit, demonstrating strong generalization across platforms.
- A significant volume of self-reported migraine content was identified on both Twitter and Reddit, confirming their utility as data sources.
- The analysis revealed diverse information on migraine therapies, including patient experiences and treatment preferences.
- Patient sentiments associated with specific therapies were consistently captured, indicating emotional and experiential dimensions of migraine management.
- The model's high performance suggests its potential for scalable, real-world monitoring of migraine-related health discourse.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.