Skip to main content
QUICK REVIEW

[論文レビュー] From Hand-Perspective Visual Information to Grasp Type Probabilities: Deep Learning via Ranking Labels

Mo Han, Sezen Yağmur Günay|arXiv (Cornell University)|Mar 8, 2021
Muscle activation and electromyography studies参考文献 32被引用数 6
ひとこと要約

本論文では、順序付きラベルを用いて握り方の確率を推定する深層学習フレームワークを提案する。このフレームワークは、握りのラベル付けにおける順列のあいまいさを扱うためにPlackett-Luceモデルを活用し、手元からの視覚入力から握り方の確率を予測する。手元の画像を学習データとし、人間が順序付けた握りリストを教師信号として用いることで、確率的かつ人間が関与する制御を実現するロボットプロステティクスの分野で、多様な握り方に対して高い耐障害性と一般化性能を達成する。

ABSTRACT

Limb deficiency severely affects the daily lives of amputees and drives efforts to provide functional robotic prosthetic hands to compensate this deprivation. Convolutional neural network-based computer vision control of the prosthetic hand has received increased attention as a method to replace or complement physiological signals due to its reliability by training visual information to predict the hand gesture. Mounting a camera into the palm of a prosthetic hand is proved to be a promising approach to collect visual data. However, the grasp type labelled from the eye and hand perspective may differ as object shapes are not always symmetric. Thus, to represent this difference in a realistic way, we employed a dataset containing synchronous images from eye- and hand- view, where the hand-perspective images are used for training while the eye-view images are only for manual labelling. Electromyogram (EMG) activity and movement kinematics data from the upper arm are also collected for multi-modal information fusion in future work. Moreover, in order to include human-in-the-loop control and combine the computer vision with physiological signal inputs, instead of making absolute positive or negative predictions, we build a novel probabilistic classifier according to the Plackett-Luce model. To predict the probability distribution over grasps, we exploit the statistical model over label rankings to solve the permutation domain problems via a maximum likelihood estimation, utilizing the manually ranked lists of grasps as a new form of label. We indicate that the proposed model is applicable to the most popular and productive convolutional neural network frameworks.

研究の動機と目的

  • プロステティクス制御のための握り認識において、目線と手元の視点の違いに起因する乖離を是正すること。
  • 握り選択の不確実性をモデル化することで、人間が関与する制御を支援する確率的握り分類手法の開発。
  • 握りのラベル付けにおける順列のあいまいさを、順序付きの握りリストを教師信号として用いることで克服すること。
  • 視覚的信号と生理的信号(EMG、運動学的データ)の融合を可能にし、マルチモーダルプロステティクス制御を実現すること。
  • 実用的導入を可能にするために、標準的なCNNアーキテクチャと互換性を持つ手法の設計。

提案手法

  • 本手法は、特徴抽出のための畳み込みニューラルネットワークに、手元からの画像を入力として用いる。
  • 視覚的認識のための教師信号として、目線からの画像における手動による握り順位付けのリストを新規に採用し、人間の握りの可能性に関する認識を表現する。
  • Plackett-Luceモデルを用いて、順序付きラベルに基づく握り方タイプの確率分布をモデル化し、最尤推定を可能にする。
  • フレームワークは、握り予測タスクを順序付けに基づく学習問題に変換することで、ラベルの順列変化に対する感受性を低減する。
  • 勾配ベースの最適化が可能となるように、Plackett-Luce尤度から導出された交差エントロピー損失を用いて、エンドツーエンドでモデルを学習する。
  • ResNet や EfficientNet といった一般的なCNNアーキテクチャと互換性があり、既存のビジョンパイプラインへの統合を可能にする。

実験結果

リサーチクエスチョン

  • RQ1物体の非対称性によって目線と手元の視点の握り方が異なる場合、どのようにして握り予測の性能を向上させられるか?
  • RQ2従来のone-hot分類と比較して、順序付きラベルを用いることで、握り予測の頑健性と現実性が向上するか?
  • RQ3Plackett-Luceモデルを用いて握りの確率をモデル化することで、マルチモーダルプロステティクス制御における性能がどの程度向上するか?
  • RQ4視覚的入力とEMG、運動学的データを効果的に統合することで、ロボットハンドのハイブリッド制御を実現するにはどうすればよいか?
  • RQ5提案手法は、実世界での導入を想定した標準的なディープラーニングフレームワークと互換性を持つのか?

主な発見

  • 提案手法は、特に形状が曖昧または非対称な物体に対して、ベースライン分類モデルと比較して高い分類精度を達成した。
  • Plackett-Luceモデルによる順序付きラベルの活用により、ラベルの順列変化に対する感受性が低下し、多様な握り方タイプにわたる一般化性能が向上した。
  • フレームワークは確率的出力を提供し、ユーザーが確率の高い握り方の順位付きリストから選択可能な人間が関与する制御を可能にした。
  • モデルは最先端のCNNアーキテクチャと互換性があり、既存のロボットビジョンパイプラインへの統合が可能である。
  • 同期された目線・手元の画像、ならびにEMGと運動学的データを含むデータセットは、将来的なマルチモーダル融合研究を支援する。
  • 物体の対称性が保証されない実世界の条件下でも、本手法はより高い頑健性を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。