Skip to main content
QUICK REVIEW

[论文解读] Generalizable machine learning for stress monitoring from wearable devices: A systematic literature review

Gideon Vos, Kelly Trinh|arXiv (Cornell University)|Sep 29, 2022
Digital Mental Health Interventions被引用 4
一句话总结

本篇系统性文献回顾评估了利用可穿戴设备进行压力监测的机器学习模型,重点关注模型在未见数据上的泛化能力。研究识别出现有研究中的关键局限性——数据集规模小、仅在单一实验环境下采集、标签不一致以及泛化能力差——并倡导建立更大、更多样化的公开数据集及标准化协议,以提升压力检测模型在现实场景中的适用性。

ABSTRACT

Introduction. The stress response has both subjective, psychological and objectively measurable, biological components. Both of them can be expressed differently from person to person, complicating the development of a generic stress measurement model. This is further compounded by the lack of large, labeled datasets that can be utilized to build machine learning models for accurately detecting periods and levels of stress. The aim of this review is to provide an overview of the current state of stress detection and monitoring using wearable devices, and where applicable, machine learning techniques utilized. Methods. This study reviewed published works contributing and/or using datasets designed for detecting stress and their associated machine learning methods, with a systematic review and meta-analysis of those that utilized wearable sensor data as stress biomarkers. The electronic databases of Google Scholar, Crossref, DOAJ and PubMed were searched for relevant articles and a total of 24 articles were identified and included in the final analysis. The reviewed works were synthesized into three categories of publicly available stress datasets, machine learning, and future research directions. Results. A wide variety of study-specific test and measurement protocols were noted in the literature. A number of public datasets were identified that are labeled for stress detection. In addition, we discuss that previous works show shortcomings in areas such as their labeling protocols, lack of statistical power, validity of stress biomarkers, and generalization ability. Conclusion. Generalization of existing machine learning models still require further study, and research in this area will continue to provide improvements as newer and more substantial datasets become available for study.

研究动机与目标

  • 评估在公开压力相关可穿戴设备数据集上训练的机器学习模型的泛化性能。
  • 识别当前压力检测研究中的方法论缺陷,包括标签协议、数据多样性及统计效能。
  • 评估用于压力预测模型的生理生物标志物(HRV、EDA、HR)的有效性与可靠性。
  • 指出缺乏对未见数据的外部验证,以及缺乏用于稳健模型训练的大规模、多样化公开数据集。
  • 通过阐明开发可泛化、适用于真实世界的压力监测系统所面临的挑战与机遇,为未来研究提供指导。

提出的方法

  • 通过 Google Scholar、Crossref、DOAJ 和 PubMed 进行系统性文献回顾,识别出 33 篇关于可穿戴设备用于压力检测的相关研究。
  • 将研究归纳为三类:公开可用的压力数据集、应用的机器学习技术,以及未来研究方向。
  • 评估模型验证方法,强调依赖留一被试者交叉验证(LOSO)或 K 折交叉验证,但缺乏对外部数据集的测试。
  • 使用 IJMEDI 检查清单评估研究质量,以确保所纳入研究的方法论严谨性。
  • 分析特征工程与模型在不同数据集上的性能表现,尤其关注 EDA 和 HR/HRV 生物标志物的纳入情况。
  • 绘制随时间变化的报告准确率,以评估模型性能与泛化能力的趋势。

实验结果

研究问题

  • RQ1现有用于压力检测的机器学习模型在未见数据和新受试者上的泛化程度如何?
  • RQ2标签协议与实验条件的差异在多大程度上影响了压力生物标志物测量的可靠性和有效性?
  • RQ3纳入特定生理生物标志物(如 EDA、HRV、HR)对模型准确率与泛化能力有何影响?
  • RQ4尽管机器学习与可穿戴技术在过去十年中持续进步,为何模型准确率并未呈现一致提升?
  • RQ5当前研究中存在哪些关键的方法论空白,阻碍了稳健、真实世界适用的压力监测系统的发展?

主要发现

  • 大多数研究使用小规模、单一实验环境下的数据集,数据量不足 24 小时,限制了统计效能与泛化潜力。
  • 仅少数研究在完全新的、未见过的数据集上验证模型,且数据采集条件不同,导致泛化能力未经验证。
  • 同时包含 EDA 和 HR(或 HRV)生物标志物的研究,准确率超过 90%;而排除任一标志物则准确率降至 86% 以下。
  • 尽管过去十年间技术与方法论持续进步,但模型准确率并未呈现一致提升,表明泛化能力存在停滞。
  • 个体特异性模型优于通用模型,表明个体水平的适应可能是实现高精度压力预测的关键。
  • 缺乏数据采集、标签与设备放置的标准化指南,仍是实现可靠且可复现压力监测的主要障碍。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。