[Paper Review] Metrics Matter in Surgical Phase Recognition
The paper analyzes how evaluation metric choices and reporting details impact comparability in surgical phase recognition on Cholec80, and provides guidance and a baseline evaluation across metric variants.
Surgical phase recognition is a basic component for different context-aware applications in computer- and robot-assisted surgery. In recent years, several methods for automatic surgical phase recognition have been proposed, showing promising results. However, a meaningful comparison of these methods is difficult due to differences in the evaluation process and incomplete reporting of evaluation details. In particular, the details of metric computation can vary widely between different studies. To raise awareness of potential inconsistencies, this paper summarizes common deviations in the evaluation of phase recognition algorithms on the Cholec80 benchmark. In addition, a structured overview of previously reported evaluation results on Cholec80 is provided, taking known differences in evaluation protocols into account. Greater attention to evaluation details could help achieve more consistent and comparable results on the surgical phase recognition task, leading to more reliable conclusions about advancements in the field and, finally, translation into clinical practice.
Motivation & Objective
- Highlight how evaluation protocol variations affect comparability of phase-recognition methods on Cholec80.
- Summarize common deviations in metric computation and reporting across studies.
- Provide a structured overview of reported results accounting for protocol differences.
- Offer recommendations to improve reproducibility and translation to clinical practice.
Proposed method
- Summarize common deviations in how metrics are computed for surgical phase recognition on Cholec80.
- Present a structured overview of existing evaluation results considering protocol differences.
- Comprehensively evaluate a baseline phase recognition model using multiple metric variants.
- Define and discuss standard video-wise and phase-wise evaluation metrics and their computation.
- Discuss strategies for handling undefined values and relaxation of timing constraints in metrics.
Experimental results
Research questions
- RQ1How do different evaluation protocols and metric implementations affect reported performance on Cholec80?
- RQ2What are the common inconsistencies in metric computation for surgical phase recognition, and how do they influence comparability?
- RQ3How does a baseline model perform across various metric variants and data-split strategies?
- RQ4What practices can improve reproducibility and fairer comparisons in surgical phase recognition literature?
Key findings
- Evaluation results on Cholec80 are not directly comparable due to variations in data splits, metric definitions, and handling of undefined values.
- Common discrepancies include relaxed versus strict boundaries, different data splits, and whether standard deviations are computed over videos or runs.
- The paper demonstrates how applying different metric variants to a baseline model yields different interpretations of performance.
- Undefined values in phase-wise metrics require explicit handling strategies, which influence overall macro- and per-phase scores.
- The authors provide a codebase for phase metrics to promote consistent evaluation across studies.
- A structured review of recent Cholec80 results highlights the challenge of drawing conclusions about state-of-the-art methods under inconsistent protocols.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.