Skip to main content
QUICK REVIEW

[論文レビュー] Towards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks

Duygu Sarıkaya, Pierre Jannin|arXiv (Cornell University)|Jan 11, 2020
Human Pose and Action Recognition参考文献 49被引用数 16
ひとこと要約

本論文は、手術用器具の関節ポーズを動的グラフとしてモデル化する新しい空間時系列グラフ畳み込みネットワーク(ST-GCN)アプローチを提案する。この手法により、低レベルの手術ジェスチャーを認識し、JIGSAWSのステーニングタスクで平均68%の精度を達成した。これは10%のランダムベースラインを著しく上回り、ポーズに基づくシーン不変表現を活用することで、異なるデータセット間での高い汎用性を示した。

ABSTRACT

Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across different tasks and different datasets remains a challenge. We introduce a modality that is robust to scene variation, and that is able to infer part information such as orientational and relative spatial relationships. The proposed modality is based on spatial temporal graph representations of surgical tools in videos, for surgical activity recognition. To explore its effectiveness, we model and recognize surgical gestures with the proposed modality. We construct spatial graphs connecting the joint pose estimations of surgical tools. Then, we connect each joint to the corresponding joint in the consecutive frames forming inter-frame edges representing the trajectory of the joint over time. We then learn hierarchical spatial temporal graph representations using Spatial Temporal Graph Convolutional Networks (ST-GCN). Our experiments show that learned spatial temporal graph representations perform well in surgical gesture recognition even when used individually. We experiment with the Suturing task of the JIGSAWS dataset where the chance baseline for gesture recognition is 10%. Our results demonstrate 68% average accuracy which suggests a significant improvement. Learned hierarchical spatial temporal graph representations can be used either individually, in cascades or as a complementary modality in surgical activity recognition, therefore provide a benchmark for future studies. To our knowledge, our paper is the first to use spatial temporal graph representations of surgical tools, and pose-based skeleton representations in general, for surgical activity recognition.

研究の動機と目的

  • 異なるタスクやデータセット間で一般化性に課題を抱える手術行動認識の課題に対処すること。
  • 手術シーン、視点、背景の変化に対して頑健な、シーン不変の表現を構築すること。
  • 手術ジェスチャー認識の補完的または独立したモodalとして、ポーズベースのスケルトン表現の利用を検討すること。
  • 空間時系列グラフ表現が、手術用器具の階層的運動および空間的関係をどれだけ効果的に捉えられるかを評価すること。
  • 学習された階層的空間時系列グラフ特徴を用いた手術動画解析分野の今後の研究のベンチマークを確立すること。

提案手法

  • フレーム間の手術用器具の関節ポーズ推定値を接続して、非有向空間グラフを構築する。
  • 連続するフレーム間の対応する関節を結ぶことで、時間的軌跡をモデル化するためのフレーム間エッジを形成する。
  • 空間時系列グラフ畳み込みネットワーク(ST-GCN)を用いて階層的空間時系列表現を学習する。
  • 深度情報やオプティカルフローに依存せずに、動画からの2次元関節座標(X, Y)を入力とする。
  • 複数層のST-GCNを適用し、進化するグラフ構造から文脈的・動的特徴を抽出する。
  • JIGSAWSステーニングデータセットに対して、1人を除いたユーザーを交差検証で学習・評価する。

実験結果

リサーチクエスチョン

  • RQ1ポーズベースの空間時系列グラフ表現は、異なるデータセット間で一般化性を向上させることができるか?
  • RQ2従来の画像ベースやキネマティクスベースのモデルと比較して、学習されたグラフベースの表現はどれほど効果的か?
  • RQ3空間時系列グラフ特徴は、手術用器具の動きにおける意味のある運動的および空間的関係をどの程度捉えることができるか?
  • RQ4提案手法のモダリティは、マルチモーダルな手術行動認識において補完的または独立した構成要素として機能できるか?
  • RQ5ST-GCNモデルの性能は、JIGSAWSベンチマークで最先端の手法と比較してどの程度か?

主な発見

  • 提案手法は、JIGSAWSのステーニングタスクで平均68%の精度を達成し、10%のランダムベースラインを著しく上回った。
  • モデルは、追加の視覚的またはキネマティクス的ヒントがなくても、単体としてのモダリティとして優れた性能を示した。
  • 空間時系列グラフ表現は、階層的運動パターンと手術用器具の関節間の相対的空間的関係を効果的に捉えた。
  • 背景や照明の変化に対して頑健であることが示された。これは、画像レベルの特徴に敏感な手法とは異なり、ポーズベースの関節情報に依存しているためである。
  • 学習された空間時系列グラフ特徴は、今後の手術行動認識研究の意味のあるベンチマークとして機能できる可能性がある。
  • 2次元関節座標のみを用いているにもかかわらず、3D CNN や ST-CNN を含む、最近の多数のCNNベースのモデルを上回る性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。