[Paper Review] CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
CoPESD introduces a fine-grained, multi-level surgical motion dataset for Endoscopic Submucosal Dissection (ESD) and demonstrates LVLMs trained on CoPESD can predict low-level robotic motions to act as an ESD co-pilot.
submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technically challenging and carries high risks of complications, necessitating skilled surgeons and precise instruments. Recent advancements in Large Visual-Language Models (LVLMs) offer promising decision support and predictive planning capabilities for robotic systems, which can augment the accuracy of ESD and reduce procedural risks. However, existing datasets for multi-level fine-grained ESD surgical motion understanding are scarce and lack detailed annotations. In this paper, we design a hierarchical decomposition of ESD motion granularity and introduce a multi-level surgical motion dataset (CoPESD) for training LVLMs as the robotic extbf{Co}- extbf{P}ilot of extbf{E}ndoscopic extbf{S}ubmucosal extbf{D}issection. CoPESD includes 17,679 images with 32,699 bounding boxes and 88,395 multi-level motions, from over 35 hours of ESD videos for both robot-assisted and conventional surgeries. CoPESD enables granular analysis of ESD motions, focusing on the complex task of submucosal dissection. Extensive experiments on the LVLMs demonstrate the effectiveness of CoPESD in training LVLMs to predict following surgical robotic motions. As the first multimodal ESD motion dataset, CoPESD supports advanced research in ESD instruction-following and surgical automation. The dataset is available at \href{https://github.com/gkw0010/CoPESD}{https://github.com/gkw0010/CoPESD.}}
Motivation & Objective
- Provide a granular, multi-level decomposition of ESD motions to enable partial automation and LVLM-based co-piloting.
- Create a publicly accessible multimodal dataset (images + bounding boxes + multi-level motions) from both robot-assisted and conventional ESD.
- Demonstrate that fine-tuned LVLMs can follow surgical instructions and predict low-level robotic motions for ESD.
- Establish benchmarks and evaluation protocols to assess instruction-following, grounding, and motion-direction accuracy in LVLMs for surgical automation.
Proposed method
- Propose a hierarchical motion granularity for ESD: operation, task, surgeme, motion primitive, and navigating motion primitive.
- Collect and process 40 ESD videos (robot-assisted and conventional) to yield 17,679 images with 32,699 bounding boxes and 88,395 multi-level motions.
- Annotate images with multi-level motions; validate annotations via cross-checks by two endoscopists and quality control; augment text variability using ChatGPT to generate five phrasing variants for each motion.
- Fine-tune state-of-the-art LVLMs (SPHINX-X and LLaVA-1.5) with LLaMA-2 backbones on CoPESD; evaluate with GPT-based response scoring, grounding (mIoU), and motion-direction accuracy/F-score.
- Evaluate ablations on image resolution and data proportion to assess impact on instruction-following and motion prediction.
Experimental results
Research questions
- RQ1Can a multi-level ESD motion dataset enable LVLMs to act as a co-pilot by predicting low-level robotic motions from visual inputs and textual prompts?
- RQ2What is the impact of image resolution and training data size on LVLMs' ability to follow ESD instructions and localize instruments?
- RQ3How do different LVLM backbones and dataset scales affect grammar-aware motion instruction generation and grounding accuracy in ESD?
- RQ4Does CoPESD support robust instruction-following and motion prediction across robot-assisted and conventional ESD scenarios?
Key findings
- LVLMs fine-tuned on CoPESD achieve higher GPT-based response quality and stronger instrument grounding than baselines.
- Higher input image resolution improves motion prediction accuracy and localization (mIoU) across models.
- Larger LLM backbones generally yield better performance across metrics and tasks.
- Using up to 100% of CoPESD yields superior motion-type and direction prediction (accuracy and F-score) compared to smaller subsets.
- CoPESD enables LVLMs to produce precise next-step motion descriptions and locate instruments in ESD scenes.
- CoPESD is the first multimodal ESD motion dataset and supports instruction-following and surgical automation research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.