[Paper Review] Robust Medical Instrument Segmentation Challenge 2019
This paper introduces the ROBUST-MIS 2019 challenge, a large-scale benchmarking effort for instrument detection and segmentation in laparoscopic videos, focusing on robustness and generalization across stages with increasing domain gaps.
Intraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts).
Motivation & Objective
- Benchmark robustness of instrument detection and segmentation in endoscopic videos.
- Assess generalization of methods across different surgeries and institutions.
- Identify image conditions that degrade performance (e.g., smoke, bleeding, motion artifacts).
- Provide a fair comparison framework through a multi-task, multi-stage challenge.
- Promote development of video-only approaches suitable for robot-assisted surgery.
Proposed method
- Organized as a MICCAI 2019 EndoVis sub-challenge with three tasks: binary segmentation, multi-instance detection, and multi-instance segmentation.
- Used a large, expert-annotated dataset of 10,040 frames from 30 procedures across three surgery types.
- Implemented three evaluation stages with increasing domain gap (Stage 1: training-patient data; Stage 2: same surgery type, different patients; Stage 3: different but similar surgery type).
- Assessed performance with DSC, NSD, MI_DSC, MI_NSD, and mAP metrics; matched instances via the Hungarian algorithm when needed.
- Implemented a two-ranking scheme (accuracy and robustness) for segmentation tasks, plus mAP for detection, including bootstrap analyses to assess ranking stability.
Experimental results
Research questions
- RQ1How do current instrument detection and segmentation methods perform on a robust, real-world surgical video dataset?
- RQ2Do state-of-the-art models generalize across different surgeries and institutions when trained on one type of procedure?
- RQ3What image artifacts or challenges most impact performance (e.g., blood, smoke, motion) across tasks?
- RQ4Does a multi-task approach (binary, multi-instance detection/segmentation) improve robustness and generalization compared to single-task methods?
- RQ5How does performance degrade under increasing domain gap, and can worst-case performance be quantified and improved?
Key findings
- Best-performing methods achieve high average accuracy, but performance degrades with larger domain gaps between training and test data.
- Robustness-focused evaluation (5th percentile) highlights worst-case limitations across tasks.
- Binary and multi-instance segmentation tasks use Dice-based and surface-based metrics (DSC, NSD) with MI_DSC/MI_NSD for per-instance assessments.
- Multi-instance detection is evaluated with mean Average Precision (mAP) using IoU threshold 0.3 for matching.
- 30 procedures and 10,040 frames provide a diverse testbed, revealing that small, crossing, moving, and transparent instrument parts remain challenging.
- The results emphasize focusing future research on detecting and segmenting small or partially visible instruments, especially under challenging image conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.