Skip to main content
QUICK REVIEW

[論文レビュー] Adaptive Feature Processing for Robust Human Activity Recognition on a Novel Multi-Modal Dataset

Mirco Moencks, De Silva|arXiv (Cornell University)|Jan 9, 2019
Human Pose and Action Recognition参考文献 68被引用数 16
ひとこと要約

本稿では、16名の参加者を対象に、RGB、深度、慣性、磁気センサーを用いた9つの屋内行動を捉えた、新規で公開可能なマルチモーダルデータセットを紹介する。また、深層ニューラルネットワークを用いてRGB-深度データで全動的および静的行動を認識する際、最大96.8%の正確性を達成する、適応的特徴処理パイプラインを提案する。

ABSTRACT

Human Activity Recognition (HAR) is a key building block of many emerging applications such as intelligent mobility, sports analytics, ambient-assisted living and human-robot interaction. With robust HAR, systems will become more human-aware, leading towards much safer and empathetic autonomous systems. While human pose detection has made significant progress with the dawn of deep convolutional neural networks (CNNs), the state-of-the-art research has almost exclusively focused on a single sensing modality, especially video. However, in safety critical applications it is imperative to utilize multiple sensor modalities for robust operation. To exploit the benefits of state-of-the-art machine learning techniques for HAR, it is extremely important to have multimodal datasets. In this paper, we present a novel, multi-modal sensor dataset that encompasses nine indoor activities, performed by 16 participants, and captured by four types of sensors that are commonly used in indoor applications and autonomous vehicles. This multimodal dataset is the first of its kind to be made openly available and can be exploited for many applications that require HAR, including sports analytics, healthcare assistance and indoor intelligent mobility. We propose a novel data preprocessing algorithm to enable adaptive feature extraction from the dataset to be utilized by different machine learning algorithms. Through rigorous experimental evaluations, this paper reviews the performance of machine learning approaches to posture recognition, and analyses the robustness of the algorithms. When performing HAR with the RGB-Depth data from our new dataset, machine learning algorithms such as a deep neural network reached a mean accuracy of up to 96.8% for classification across all stationary and dynamic activities

研究の動機と目的

  • 安全上の重要な応用分野における、公開可能でマルチモーダルなHARデータセットの不足に対処すること。
  • 多様なセンサーモダリティからの特徴抽出を可能にする、適応的データ前処理パイプラインの開発。
  • 複数のセンサータイプを用いたポーズおよび行動認識のための機械学習モデルの評価と比較。
  • 複数のセンシングモダリティからの補完的情報を活用することで、システムの耐障害性を向上させること。
  • 今後のHAR分野の研究、特に医療、移動支援、ロボット工学分野におけるベンチマークデータセットと手法の提供。

提案手法

  • 著者らは、4種類のセンサータイプ(RGBカメラ、深度センサー、慣性計測単位(IMU)、磁気計)を用いてマルチモーダルデータセットを収集した。
  • 信号の特性と時間的ダイナミクスに基づき、各モダリティからの特徴抽出および正規化を実行する適応的特徴処理アルゴリズムを設計した。
  • 入力表現を標準化することで、深層ニューラルネットワーク(DNN)を含むさまざまな機械学習モデルとの柔軟な統合を可能とした。
  • 特徴抽出には、時間領域統計、周波数領域変換(例:FFT)、およびRGBおよび深度ストリームのための空間的時間的符号化が含まれる。
  • センサータイプおよび行動クラスに応じて前処理を動的に調整することで、モデルの汎化性能を向上させた。
  • 最終分類のため、モダリティ固有の特徴を統合するマルチストリームDNNアーキテクチャを用いた。

実験結果

リサーチクエスチョン

  • RQ1提案された適応的前処理を用いた場合、HARモデルの性能は、異なるセンサーモダリティでどのように変化するか?
  • RQ2単一モダリティ手法と比較して、マルチモーダル統合は、認識正確性と耐障害性をどの程度向上させるか?
  • RQ3提案された適応的特徴処理パイプラインは、多様な人間の行動にわたるモデルの汎化性能を向上させるのにどの程度有効か?
  • RQ4最新のSOTAモデルがこの新規マルチモーダルデータセットで達成できる性能の上限は何か?
  • RQ5参加者および行動の多様性が、モデルの転送性および耐障害性にどのように影響するか?

主な発見

  • 提案された適応的特徴処理パイプラインは、ベースライン前処理と比較して、すべてのセンサーモダリティでモデル性能を顕著に向上させた。
  • 深層ニューラルネットワークは、新規データセットのRGB-深度データを用いることで、平均分類正確性96.8%を達成した。
  • マルチモーダル統合は単一モダリティモデルを常に上回り、特にRGB-深度の組み合わせが最高の正確性を示した。
  • このデータセットは、異なる参加者および行動タイプにおいて一貫した性能を示し、強力な汎化可能性を示した。
  • 慣性および磁気センサーの統合により、視認性が低いまたは遮蔽された状況でも耐障害性が向上した。
  • 本データセットは、4種類のセンサータイプを備えた、公開可能な最初のマルチモーダルHARデータセットであり、人間中心のAI応用分野における広範な研究を可能にする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。