[Paper Review] WellXplain: Wellness Concept Extraction and Classification in Reddit Posts for Mental Health Analysis
WellXplain introduces a novel dataset of 3,092 Reddit posts annotated for 7 wellness dimensions based on Halbert L. Dunn’s theory, enabling fine-grained mental health screening through concept extraction and explainable classification. The study establishes baselines using diverse models, demonstrating that domain-adapted transformers and LLMs outperform traditional classifiers, with human-annotated spans enhancing model interpretability.
During the current mental health crisis, the importance of identifying potential indicators of mental issues from social media content has surged. Overlooking the multifaceted nature of mental and social well-being can have detrimental effects on one's mental state. In traditional therapy sessions, professionals manually pinpoint the origins and outcomes of underlying mental challenges, a process both detailed and time-intensive. We introduce an approach to this intricate mental health analysis by framing the identification of wellness dimensions in Reddit content as a wellness concept extraction and categorization challenge. We've curated a unique dataset named WELLXPLAIN, comprising 3,092 entries and totaling 72,813 words. Drawing from Halbert L. Dunn's well-regarded wellness theory, our team formulated an annotation framework along with guidelines. This dataset also includes human-marked textual segments, offering clear reasoning for decisions made in the wellness concept categorization process. Our aim in publishing this dataset and analyzing initial benchmarks is to spearhead the creation of advanced language models tailored for healthcare-focused concept extraction and categorization.
Motivation & Objective
- Address the gap in multi-dimensional mental health screening by identifying wellness dimensions in social media text, moving beyond symptom-focused analysis.
- Develop a reliable, explainable framework for detecting early signs of mental distress through contextual, clinically grounded wellness concepts.
- Create a high-quality, human-annotated dataset with explicit text spans to support model interpretability and clinical validation.
- Enable longitudinal mental health monitoring by capturing shifts in wellness dimensions over time in user-generated content.
- Facilitate the development of robust, explainable AI models for mental health applications using domain-specific and large language models.
Proposed method
- Construct a new dataset, WellXplain, comprising 3,092 Reddit posts with 72,813 words, annotated for 7 wellness dimensions derived from Halbert L. Dunn’s theory of wellness.
- Design a detailed annotation scheme and perplexity guidelines to ensure consistency and clinical relevance in labeling wellness concepts.
- Collect human-annotated text spans as explanations for each wellness concept prediction to enhance model interpretability and support human-in-the-loop validation.
- Train and evaluate diverse models, including traditional multi-class classifiers, domain-adapted BERT variants, and large language models (LLMs), on the wellness concept classification task.
- Use the annotated spans as auxiliary supervision to guide model attention and improve focus on relevant linguistic cues in the text.
- Evaluate model performance using standard metrics (e.g., F1, accuracy) and compare generalization across model types, emphasizing reliability and explainability.
Experimental results
Research questions
- RQ1Can a multi-dimensional wellness concept extraction framework improve early detection of mental health concerns in social media text compared to symptom-only approaches?
- RQ2How do domain-adapted transformers and large language models perform in classifying wellness dimensions compared to traditional classifiers on Reddit posts?
- RQ3To what extent do human-annotated explanation spans improve model performance and interpretability in wellness concept classification?
- RQ4How stable and reliable are model predictions across different wellness dimensions, and what are the key challenges in detecting subtle shifts in well-being over time?
- RQ5Can the WellXplain dataset support longitudinal mental health analysis by capturing evolving wellness states in user posts?
Key findings
- Domain-adapted transformers and large language models outperformed traditional multi-class classifiers in classifying wellness dimensions, indicating their suitability for nuanced mental health analysis.
- The inclusion of human-annotated explanation spans significantly improved model attention and interpretability, supporting human validation and trust in predictions.
- Despite strong performance, model predictions remained insufficient for clinical diagnosis, underscoring the need for human oversight and ethical safeguards.
- The dataset revealed that wellness dimensions such as emotional, social, and spiritual well-being often shift in tandem with political and relational stressors, reflecting complex mental health trajectories.
- Longitudinal analysis of user posts showed progressive deterioration in wellness dimensions—e.g., from confusion (T1) to suicidal ideation (T5)—highlighting the value of temporal monitoring.
- The WellXplain dataset demonstrated potential for use in policy and clinical research by enabling holistic, multi-dimensional evaluation of well-being aligned with SDG 3.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.