Skip to main content
QUICK REVIEW

[论文解读] Human Activity Recognition Using Tools of Convolutional Neural Networks: A State of the Art Review, Data Sets, Challenges and Future Prospects

Md. Milon Islam, Sheikh Nooruddin|arXiv (Cornell University)|Feb 2, 2022
Context-Aware Activity Recognition Systems被引用 10
一句话总结

本文对基于卷积神经网络(CNN)的人体活动识别(HAR)技术进行了全面的最新进展综述,分析了智能手机、雷达、视觉及多模态传感器系统。评估了CNN架构在性能、超参数和数据集方面的表现,识别出关键挑战,并提出了HAR应用的未来研究方向。

ABSTRACT

Human Activity Recognition (HAR) plays a significant role in the everyday life of people because of its ability to learn extensive high-level information about human activity from wearable or stationary devices. A substantial amount of research has been conducted on HAR and numerous approaches based on deep learning and machine learning have been exploited by the research community to classify human activities. The main goal of this review is to summarize recent works based on a wide range of deep neural networks architecture, namely convolutional neural networks (CNNs) for human activity recognition. The reviewed systems are clustered into four categories depending on the use of input devices like multimodal sensing devices, smartphones, radar, and vision devices. This review describes the performances, strengths, weaknesses, and the used hyperparameters of CNN architectures for each reviewed system with an overview of available public data sources. In addition, a discussion with the current challenges to CNN-based HAR systems is presented. Finally, this review is concluded with some potential future directions that would be of great assistance for the researchers who would like to contribute to this field.

研究动机与目标

  • 提供对基于CNN的人体活动识别(HAR)方法在多种输入模态下的系统性综述。
  • 分析并比较不同用于HAR的CNN架构在性能、优势、劣势及超参数方面的表现。
  • 整理并评估当前研究中使用的公开可用HAR数据集。
  • 识别并讨论基于CNN的HAR系统中持续存在的挑战,如泛化能力、数据不平衡及实时部署问题。
  • 提出未来研究方向,以推动利用深度学习实现更鲁棒、高效且可扩展的HAR解决方案。

提出的方法

  • 将HAR系统按四种输入模态分类:多模态传感设备、智能手机、雷达及基于视觉的系统。
  • 系统分析各类别中使用的CNN架构,包括网络深度、滤波器大小、激活函数及优化技术。
  • 使用标准指标(如准确率、F1值及精确率/召回率)在基准数据集上评估模型性能。
  • 整理并描述14个公开可用的HAR数据集,包括传感器类型、采样率及活动类别。
  • 识别在所综述文献中常见的超参数配置(例如学习率、批量大小)。
  • 综合分析实际部署中的挑战,包括领域偏移、传感器差异性及模型可解释性问题。

实验结果

研究问题

  • RQ1在不同HAR输入模态中,哪些CNN架构展现出最高的准确率和鲁棒性?
  • RQ2超参数选择(如学习率、批量大小)如何影响CNN在HAR任务中的性能?
  • RQ3最广泛使用且公开可用的HAR数据集有哪些?其特性如何影响模型性能?
  • RQ4阻碍基于CNN的HAR系统在真实环境中部署的主要技术和实际挑战是什么?
  • RQ5未来哪些研究方向最有可能提升基于CNN的HAR模型的泛化能力、效率及可解释性?

主要发现

  • 基于CNN的HAR系统在UCL和UniMIB等基准数据集上准确率超过95%,其中深层网络如ResNet和DenseNet优于简单模型。
  • 基于智能手机的HAR系统因加速度计和陀螺仪提供的丰富高分辨率传感器数据,持续表现出高精度(95–98%准确率)。
  • 基于雷达的HAR系统在保护隐私的环境中展现出潜力,但在复杂运动条件下仍面临准确率挑战。
  • 基于视觉的HAR系统可实现高达97%的准确率,但需要大量计算资源,且对光照条件和遮挡敏感。
  • 融合惯性传感器与视觉数据的多模态融合方法可提升鲁棒性与准确率,尤其在复杂场景中表现更优。
  • 尽管在标准数据集上表现优异,但所有模态在泛化到未见个体和环境方面仍存在显著局限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。