[Paper Review] Generalizable machine learning for stress monitoring from wearable devices: A systematic literature review
This systematic literature review evaluates machine learning models for stress monitoring using wearable devices, focusing on model generalization across unseen data. It identifies critical limitations in existing studies—small, single-setting datasets, inconsistent labeling, and poor generalization—while advocating for larger, diverse public datasets and standardized protocols to improve real-world applicability of stress detection models.
Introduction. The stress response has both subjective, psychological and objectively measurable, biological components. Both of them can be expressed differently from person to person, complicating the development of a generic stress measurement model. This is further compounded by the lack of large, labeled datasets that can be utilized to build machine learning models for accurately detecting periods and levels of stress. The aim of this review is to provide an overview of the current state of stress detection and monitoring using wearable devices, and where applicable, machine learning techniques utilized. Methods. This study reviewed published works contributing and/or using datasets designed for detecting stress and their associated machine learning methods, with a systematic review and meta-analysis of those that utilized wearable sensor data as stress biomarkers. The electronic databases of Google Scholar, Crossref, DOAJ and PubMed were searched for relevant articles and a total of 24 articles were identified and included in the final analysis. The reviewed works were synthesized into three categories of publicly available stress datasets, machine learning, and future research directions. Results. A wide variety of study-specific test and measurement protocols were noted in the literature. A number of public datasets were identified that are labeled for stress detection. In addition, we discuss that previous works show shortcomings in areas such as their labeling protocols, lack of statistical power, validity of stress biomarkers, and generalization ability. Conclusion. Generalization of existing machine learning models still require further study, and research in this area will continue to provide improvements as newer and more substantial datasets become available for study.
Motivation & Objective
- To assess the generalization performance of machine learning models trained on public stress-related wearable datasets.
- To identify methodological shortcomings in current stress detection studies, including labeling protocols, data diversity, and statistical power.
- To evaluate the validity and reliability of physiological biomarkers (HRV, EDA, HR) used in stress prediction models.
- To highlight the lack of external validation on unseen data and the absence of large, diverse public datasets for robust model training.
- To guide future research by outlining challenges and opportunities in developing generalizable, real-world applicable stress monitoring systems.
Proposed method
- Conducted a systematic literature review using Google Scholar, Crossref, DOAJ, and PubMed to identify 33 relevant studies on stress detection using wearable devices.
- Synthesized studies into three categories: publicly available stress datasets, machine learning techniques applied, and future research directions.
- Evaluated model validation approaches, emphasizing reliance on leave-one-subject-out (LOSO) or K-fold cross-validation without external dataset testing.
- Assessed study quality using the IJMEDI checklist to ensure methodological rigor in included works.
- Analyzed feature engineering and model performance across datasets, particularly focusing on inclusion of EDA and HR/HRV biomarkers.
- Mapped reported accuracy rates over time to assess trends in model performance and generalization capability.
Experimental results
Research questions
- RQ1To what extent do existing machine learning models for stress detection generalize to unseen data and new participants?
- RQ2How do variations in labeling protocols and experimental conditions affect the reliability and validity of stress biomarker measurements?
- RQ3What is the impact of including specific physiological biomarkers (e.g., EDA, HRV, HR) on model accuracy and generalization?
- RQ4Why has there been no consistent improvement in model accuracy over time despite advances in machine learning and wearable technology?
- RQ5What are the key methodological gaps in current research that hinder the development of robust, real-world stress monitoring systems?
Key findings
- Most reviewed studies used small, single-experimental-setting datasets with less than 24 hours of data, limiting statistical power and generalization potential.
- Only a minority of studies validated models on completely new, unseen datasets collected under different conditions, leaving generalization unverified.
- Studies that included both EDA and HR (or HRV) biomarkers reported accuracy rates above 90%, whereas excluding either biomarker reduced accuracy to below 86%.
- Despite technological and methodological advances over the past decade, no consistent improvement in model accuracy was observed over time, indicating stagnation in generalization capability.
- Person-specific models outperformed generic models, suggesting individual-level adaptation may be essential for high-accuracy stress prediction.
- A lack of standardized guidelines for data collection, labeling, and device placement remains a major barrier to reliable and reproducible stress monitoring.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.