Skip to main content
QUICK REVIEW

[論文レビュー] Identifying Emotions from Walking using Affective and Deep Features

Tanmay Randhavane, Uttaran Bhattacharya|arXiv (Cornell University)|Jun 14, 2019
Human Pose and Action Recognition参考文献 86被引用数 50
ひとこと要約

この論文は、LSTM由来の深層歩容特徴と3D歩容データから作成されたハンドクラフト型感情特徴を組み合わせたデータ駆動モデルを提示し、歩行動画から知覚感情(幸せ、悲しみ、怒り、中立)を分類する際のRandom Forest分類器で80.07%の精度を達成します。

ABSTRACT

We present a new data-driven model and algorithm to identify the perceived emotions of individuals based on their walking styles. Given an RGB video of an individual walking, we extract his/her walking gait in the form of a series of 3D poses. Our goal is to exploit the gait features to classify the emotional state of the human into one of four emotions: happy, sad, angry, or neutral. Our perceived emotion recognition approach uses deep features learned via LSTM on labeled emotion datasets. Furthermore, we combine these features with affective features computed from gaits using posture and movement cues. These features are classified using a Random Forest Classifier. We show that our mapping between the combined feature space and the perceived emotional state provides 80.07% accuracy in identifying the perceived emotions. In addition to classifying discrete categories of emotions, our algorithm also predicts the values of perceived valence and arousal from gaits. We also present an EWalk (Emotion Walk) dataset that consists of videos of walking individuals with gaits and labeled emotions. To the best of our knowledge, this is the first gait-based model to identify perceived emotions from videos of walking individuals.

研究の動機と目的

  • 歩行からの感情自動認識を社会的認知とHCI文脈における非言語的手がかりとして動機づける。
  • 深層時系列特徴とハンドクラフト型感情特徴の両方を用いる歩容ベースの感情認識パイプラインを開発する。
  • 研究用途のために歩容ビデオと知覚感情ラベルを含むEWalkデータセットを作成・公開する。

提案手法

  • RGB動画からTimePoseNetを用いて3D歩容ポーズを抽出し、フレーム間の16関節の系列を取得する。
  • 歩容系列から姿勢・動作などの感情特徴を計算する。体積、面積、距離、角度、歩幅、動作の大きさ(速度、加速度、ジャーク)を含む。
  • 複数の歩容データセットで訓練したLSTMネットワークを用いて、歩容の時間的ダイナミクスを捉える。
  • 正規化した深層特徴と感情特徴を結合し、4つの知覚感情(happy、angry、sad、neutral)のいずれかを予測するRandom Forest分類器を訓練する。
  • LSTMを複数の歩容データセットで訓練し、最終分類には最大深さ5の10-estimator RFを使用する。
  • 1384の歩容データに対して10-foldクロスバリデーションで評価し、従来の歩容ベース手法に対する精度の改善を報告する。

実験結果

リサーチクエスチョン

  • RQ1組み合わせた深層歩容特徴と感情的な姿勢/動作特徴を用いることで、歩行からの知覚感情認識の精度は、いずれかの特徴タイプだけを用いた場合と比べて改善されるか?
  • RQ2融合特徴空間を離散的な知覚感情へマッピングするのに最適な分類器は何か?
  • RQ3 actedinデータセットと非演技データセットの一般化性能はどの程度か?
  • RQ4歩容から価値性/覚醒を知覚感情のカテゴリに加えて予測できるか?
  • RQ5新しい公開歩容データセット(EWalk)を知覚感情ラベル付きで導入することが性能に与える影響は何か?

主な発見

  • 提案手法は、歩容データから4つの知覚感情カテゴリを分類する際に80.07%の精度を達成する。
  • 正規化された結合特徴(深層+感情特徴)を用いたRFは、このタスクでSVM系より優れている。
  • 非演技データセット(CMUおよびICT)でも高い性能を維持し、79.72%の精度を達成。
  • 機械的なクラウドソーシングを通じて収集された知覚感情ラベル付きの新規データセット、EWalkが導入される。
  • 彼らの評価設定で他の歩容ベース感情認識手法より13.85%の精度改善を報告。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。