Skip to main content
QUICK REVIEW

[Paper Review] A comprehensive survey on recent deep learning-based methods applied to surgical data

Mansoor Ali, Rafael Martinez Garcia Pena|arXiv (Cornell University)|Sep 3, 2022
Surgical Simulation and Training4 citations
TL;DR

This paper presents a comprehensive survey of deep learning-based methods for surgical data analysis, focusing on surgical tool localization, segmentation, tracking, and 3D scene perception in minimally invasive surgery. It reviews state-of-the-art techniques, evaluates publicly available benchmark datasets, and identifies key challenges such as limited annotated data, artifacts in endoscopic images, and non-textured tissues, while proposing directions for improving clinical translation and model robustness.

ABSTRACT

Minimally invasive surgery is highly operator dependant with a lengthy procedural time causing fatigue to surgeon and risks to patients such as injury to organs, infection, bleeding, and complications of anesthesia. To mitigate such risks, real-time systems are desired to be developed that can provide intra-operative guidance to surgeons. For example, an automated system for tool localization, tool (or tissue) tracking, and depth estimation can enable a clear understanding of surgical scenes preventing miscalculations during surgical procedures. In this work, we present a systematic review of recent machine learning-based approaches including surgical tool localization, segmentation, tracking, and 3D scene perception. Furthermore, we provide a detailed overview of publicly available benchmark datasets widely used for surgical navigation tasks. While recent deep learning architectures have shown promising results, there are still several open research problems such as a lack of annotated datasets, the presence of artifacts in surgical scenes, and non-textured surfaces that hinder 3D reconstruction of the anatomical structures. Based on our comprehensive review, we present a discussion on current gaps and needed steps to improve the adaptation of technology in surgery.

Motivation & Objective

  • To systematically review recent deep learning-based approaches for surgical navigation tasks including tool localization, segmentation, tracking, and 3D scene reconstruction.
  • To identify and analyze key challenges in surgical data science, such as lack of annotated datasets, image artifacts, and non-textured anatomical surfaces.
  • To evaluate publicly available benchmark datasets used in surgical navigation research and assess their diversity and utility.
  • To discuss open research problems and propose actionable steps for improving model generalizability, real-time performance, and clinical adoption.
  • To explore the role of emerging technologies like federated learning, self-supervised learning, and augmented reality in overcoming data and deployment limitations.

Proposed method

  • Conducting a systematic literature review of deep learning applications in surgical data science from 2016 to 2021.
  • Categorizing and analyzing methods based on their application in surgical tool detection, segmentation, tracking, and 3D reconstruction.
  • Evaluating the performance and limitations of state-of-the-art models using metrics such as mAP, Dice score, and inference speed.
  • Surveying and summarizing 15+ publicly available surgical datasets, including their annotation protocols, modalities, and task-specific applications.
  • Applying qualitative synthesis to identify trends in model architectures (e.g., U-Net, YOLO, Transformers) and training strategies (e.g., self-supervised learning, adversarial data augmentation).
  • Proposing a framework for clinical deployment that emphasizes real-time inference, hardware compatibility, and integration with pre-operative imaging and robotic systems.

Experimental results

Research questions

  • RQ1What are the most prominent deep learning architectures and techniques used in surgical tool segmentation and tracking?
  • RQ2How do current methods handle challenges such as image artifacts, low texture, and occlusions in endoscopic videos?
  • RQ3What are the key limitations of existing benchmark datasets in terms of diversity, scale, and annotation quality?
  • RQ4To what extent do current methods support real-time inference and clinical deployment in operating rooms?
  • RQ5What strategies—such as federated learning or self-supervised learning—can mitigate data scarcity and privacy concerns in surgical AI?

Key findings

  • Surgical tool segmentation and detection have received the most research attention, with U-Net and YOLO-based models achieving high Dice scores (up to 0.92) on standard benchmarks.
  • Self-supervised and attention-based methods have shown improved robustness to artifacts and low-contrast conditions in endoscopic images.
  • Despite progress, only a small fraction of models are evaluated for real-time inference, with most requiring >50 ms per frame, limiting clinical usability.
  • Publicly available datasets such as EndoCV, Cholec80, and JIGSAW are widely used but suffer from limited diversity in surgical specialties and instrument types.
  • The integration of pre-operative imaging with intra-operative deep learning models remains underexplored, despite its potential for enhancing surgical navigation.
  • Federated learning and synthetic data generation are emerging as promising solutions to address data scarcity and privacy issues in multi-center surgical AI development.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.