Skip to main content
QUICK REVIEW

[論文レビュー] Statistical and Spatio-temporal Hand Gesture Features for Sign Language Recognition using the Leap Motion Sensor

Jordan J. Bird|arXiv (Cornell University)|Feb 22, 2022
Hand Gesture Recognition Systems被引用数 7
ひとこと要約

本研究では、Leap Motionセンサーを用いた手話認識のためのハイブリッド特徴抽出手法を提案し、統計的特徴と空間時間的特徴を組み合わせることで、ジェスチャー分類の精度を向上させた。最も優れたモデルは、240個の統計的特徴と240個の空間時間的特徴を用い、86.75%の精度を達成した。これは、動的で手話認識において、特徴の融合が単一特徴アプローチを上回ることを示している。

ABSTRACT

In modern society, people should not be identified based on their disability, rather, it is environments that can disable people with impairments. Improvements to automatic Sign Language Recognition (SLR) will lead to more enabling environments via digital technology. Many state-of-the-art approaches to SLR focus on the classification of static hand gestures, but communication is a temporal activity, which is reflected by many of the dynamic gestures present. Given this, temporal information during the delivery of a gesture is not often considered within SLR. The experiments in this work consider the problem of SL gesture recognition regarding how dynamic gestures change during their delivery, and this study aims to explore how single types of features as well as mixed features affect the classification ability of a machine learning model. 18 common gestures recorded via a Leap Motion Controller sensor provide a complex classification problem. Two sets of features are extracted from a 0.6 second time window, statistical descriptors and spatio-temporal attributes. Features from each set are compared by their ANOVA F-Scores and p-values, arranged into bins grown by 10 features per step to a limit of the 250 highest-ranked features. Results show that the best statistical model selected 240 features and scored 85.96% accuracy, the best spatio-temporal model selected 230 features and scored 80.98%, and the best mixed-feature model selected 240 features from each set leading to a classification accuracy of 86.75%. When all three sets of results are compared (146 individual machine learning models), the overall distribution shows that the minimum results are increased when inputs are any number of mixed features compared to any number of either of the two single sets of features.

研究の動機と目的

  • 統計的特徴と空間時間的手のジェスチャー特徴が、手話認識の精度に与える影響を調査すること。
  • 統計的特徴のみ、空間時間的特徴のみ、または両方の組み合わせを用いたモデルの分類性能を比較すること。
  • ANOVA Fスコアとp値を用いて、各セットからの最適な特徴数を同定し、モデルの汎化性能を向上させること。
  • 特徴の融合が、単一特徴アプローチに比べ、動的ジェスチャー認識における頑健性と精度を向上させるかどうかを評価すること。

提案手法

  • 統計的特徴は、手のジェスチャーデータの0.6秒間の時間窓から抽出され、位置と速度の平均、分散、および高次モーメントを捉えた。
  • 空間時間的特徴は、軌道ダイナミクスから導出され、加速度、曲率、および時間的進行パターンを含む。
  • 特徴選択は、ANOVA Fスコアとp値を用い、特徴をランク付けし、有意性に基づいて各セットから最大250個を選択した。
  • 選択された特徴の組み合わせを用いて、合計で146体の機械学習モデルを訓練した。一般化性能が優れていることから、ランダムフォレストをベース分類器とした。
  • モデルは10分割交差検証を用いて評価され、精度、適合率、再現率、F1スコアなどの指標が使用された。
  • 特徴の融合は、初期融合によって実装され、統計的特徴と空間時間的特徴がモデル学習の前に連結された。

実験結果

リサーチクエスチョン

  • RQ1Leap Motionデータを用いた動的手話ジェスチャーの分類において、統計的特徴と空間時間的特徴はどのように比較されるか?
  • RQ2分類精度を最大化するために、各カテゴリ(統計的特徴と空間時間的特徴)からの最適な特徴数は何か?
  • RQ3統計的特徴と空間時間的特徴を組み合わせることで、単独で使用する場合よりも優れた性能が得られるか?
  • RQ4特徴の融合は、交差検証スコアの標準偏差で測定されるモデルの安定性と頑健性にどのように影響するか?
  • RQ5ANOVA Fスコアに基づく特徴選択は、モデル性能と計算複雑度にどのような影響を与えるか?

主な発見

  • 最も優れたモデルは、240個の統計的特徴と240個の空間時間的特徴を用い、平均精度86.75%、F1スコア0.867を達成した。
  • 統計的特徴のみを用いたモデルは85.96%の精度を達成し、純粋な空間時間的特徴モデル(80.98%)を上回った。これは、このデータセットにおいて統計的特徴がより判別力があることを示している。
  • 特徴を混合した場合、単一の特徴タイプのみを用いた最悪のモデルでさえ、優れた性能を示した。これは、融合により頑健性が向上したことを示している。
  • 上位10のモデルのうち7つがハイブリッドモデルであり、8番目に優れたモデルも統計的特徴のみのモデルであった。これは、混合特徴が一貫して優れた結果をもたらすことを示している。
  • 最良のハイブリッドモデルの標準偏差(0.90)は、最良の統計的特徴のみのモデル(0.51)よりも高かった。これは、特徴タイプを組み合わせることで、精度と安定性のトレードオフが生じることを示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。