Skip to main content
QUICK REVIEW

[論文レビュー] CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

Guankun Wang, Han Xiao|arXiv (Cornell University)|Oct 10, 2024
Gastric Cancer Management and Outcomes被引用数 5
ひとこと要約

CoPESD は Endoscopic Submucosal Dissection (ESD) のための細粒度かつ多レベルの手術動作データセットを導入し、CoPESD で訓練された LVLM が低レベルのロボット動作を予測して ESD のコ-pilot を務けることを示す。

ABSTRACT

submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technically challenging and carries high risks of complications, necessitating skilled surgeons and precise instruments. Recent advancements in Large Visual-Language Models (LVLMs) offer promising decision support and predictive planning capabilities for robotic systems, which can augment the accuracy of ESD and reduce procedural risks. However, existing datasets for multi-level fine-grained ESD surgical motion understanding are scarce and lack detailed annotations. In this paper, we design a hierarchical decomposition of ESD motion granularity and introduce a multi-level surgical motion dataset (CoPESD) for training LVLMs as the robotic extbf{Co}- extbf{P}ilot of extbf{E}ndoscopic extbf{S}ubmucosal extbf{D}issection. CoPESD includes 17,679 images with 32,699 bounding boxes and 88,395 multi-level motions, from over 35 hours of ESD videos for both robot-assisted and conventional surgeries. CoPESD enables granular analysis of ESD motions, focusing on the complex task of submucosal dissection. Extensive experiments on the LVLMs demonstrate the effectiveness of CoPESD in training LVLMs to predict following surgical robotic motions. As the first multimodal ESD motion dataset, CoPESD supports advanced research in ESD instruction-following and surgical automation. The dataset is available at \href{https://github.com/gkw0010/CoPESD}{https://github.com/gkw0010/CoPESD.}}

研究の動機と目的

  • ESD 動作の粒度の高い多レベル分解を提供して部分自動化と LVLM ベースのコ-piloting を可能にする。
  • ロボット支援および従来の ESD の両方からの画像 + バウンディングボックス + 多レベル動作を含む公開可能なマルチモーダルデータセットを作成する。
  • ファインチューニング済み LVLM が外科手術指示に従い、ESD の低レベルロボット動作を予測できることを実証する。
  • LVLM の外科手術自動化における指示追従、グラウンディング、動作方向性の精度を評価するベンチマークと評価プロトコルを確立する。

提案手法

  • ESD の階層的な動作粒度を提案する: operation, task, surgeme, motion primitive, and navigating motion primitive.
  • ロボット支援および従来型の 40 本の ESD ビデオを収集・処理して、32,699 のバウンディングボックスと 88,395 の多レベル動作を含む 17,679 枚の画像を得る。
  • 多レベル動作で画像に注釈を付ける;二名の内視鏡医によるクロスチェックと品質管理で注釈を検証;各動作につき five phrasing variants を生成するために ChatGPT を用いてテキスト多様性を増強。
  • CoPESD 上で最先端 LVLMs(SPHINX-X および LLaVA-1.5)を LLaMA-2 バックボーンでファインチューニング;GPT ベースの応答スコアリング、 grounding(mIoU)、および動作方向性の精度/ F スコアで評価。
  • 指示追従と動作予測への影響を評価するために画像解像度とデータ割合のアブレーションを実施。

実験結果

リサーチクエスチョン

  • RQ1ビジュアル入力とテキストプロンプトから低レベルのロボット動作を予測することで LVLM がコ-pilot として機能する多レベル ESD 動作データセットを可能にするか?
  • RQ2画像解像度と訓練データ量が LVLM の ESD 指示追従能力と器具の局在化に与える影響は何か?
  • RQ3異なる LVLM バックボーンとデータセット規模が、ESD における文法意識のある動作指示生成とグラウンディング精度にどのような影響を及ぼすか?
  • RQ4CoPESD はロボット支援および従来型 ESD シナリオ全体で堅牢な指示追従と動作予測をサポートするか?

主な発見

  • CoPESD でファインチューニングされた LVLM は、ベースラインより高い GPT ベースの応答品質とより強い器具グラウンディングを達成する。
  • 入力画像解像度が高いほど、モデル間で動作予測精度と局在化 (mIoU) が向上する。
  • より大きい LLM バックボーンは、一般に指標とタスク全体でより良い性能を示す。
  • CoPESD の最大 100% を使用することで、動作タイプと方向予測の精度と F スコアがより小さなサブセットより優れている。
  • CoPESD は LVLM が ESD シーンで正確な次ステップの動作説明を生成し、器具を特定できるようにする。
  • CoPESD は初のマルチモーダル ESD 動作データセットであり、指示追従と外科自動化研究を支援する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。