[论文解读] Comparative Validation of Machine Learning Algorithms for Surgical Workflow and Skill Analysis with the HeiChole Benchmark
本研究使用HeiChole基准数据集——一个包含33例腹腔镜胆囊切除术视频的多中心数据集——评估了用于外科手术流程与技能分析的机器学习算法。尽管在阶段识别方面表现良好(F1最高达67.7%),但动作识别与技能评估仍具挑战性,F1分数仅为21.8%–23.3%,平均误差为0.78,表明尽管潜力可观,该领域尚未完全解决。
PURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center dataset. In this work we investigated the generalizability of phase recognition algorithms in a multi-center setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 hours was created. Labels included annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 teams submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n=9 teams), for instrument presence detection between 38.5% and 63.8% (n=8 teams), but for action recognition only between 21.8% and 23.3% (n=5 teams). The average absolute error for skill assessment was 0.78 (n=1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but are not solved yet, as shown by our comparison of algorithms. This novel benchmark can be used for comparable evaluation and validation of future work.
研究动机与目标
- 评估外科阶段识别算法在单中心数据集之外的泛化能力。
- 在多中心环境下评估机器学习模型在手术动作识别、器械检测与技能评估方面的表现。
- 建立标准化基准,用于外科人工智能系统的对比评估。
- 识别当前外科手术流程与技能分析算法中的性能差距。
提出的方法
- 从三个外科中心收集了33例腹腔镜胆囊切除术视频的多中心数据集,总手术时长达22小时。
- 标注内容包括七个外科阶段(250次转换)、5,514个手术动作、21种类型中的6,980个器械实例,以及五个维度上的495次技能评估。
- 该数据集被用于2019年内窥镜视觉挑战赛的子挑战——手术流程与技能分析,共有12支团队提交了算法。
- 通过阶段、动作和器械识别的F1分数,以及技能评估的平均绝对误差对算法进行评估。
- 在各团队之间进行性能比较,以评估其在真实外科环境中的鲁棒性与泛化能力。
- 确立HeiChole基准作为未来研究的标准化评估平台。
实验结果
研究问题
- RQ1外科阶段识别的机器学习模型在多个外科中心之间的泛化能力如何?
- RQ2与阶段和器械识别相比,模型在识别手术动作方面的表现如何?
- RQ3机器学习模型能否在多个维度上准确评估外科技能?
- RQ4当前外科手术流程与技能分析系统的主要局限性是什么?
- RQ5HeiChole基准如何实现对外科人工智能算法的公平且可比的评估?
主要发现
- 九支团队的阶段识别F1分数在23.9%至67.7%之间,表明性能差异显著,仍有改进空间。
- 八支团队的器械存在性检测F1分数在38.5%至63.8%之间,表现中等但不一致。
- 五支团队的动作识别性能最弱,F1分数在21.8%至23.3%之间。
- 技能评估的平均绝对误差为0.78,表明定量技能评估存在显著不准确性。
- HeiChole基准表明,当前算法尚不足以在真实外科环境中实现稳健与可靠的应用。
- 本研究证实,尽管技术潜力可观,外科手术流程与技能分析仍是未解决的挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。