Skip to main content
QUICK REVIEW

[論文レビュー] Comparative Validation of Machine Learning Algorithms for Surgical Workflow and Skill Analysis with the HeiChole Benchmark

Martin Wagner, Beat P. Müller‐Stich|arXiv (Cornell University)|Sep 30, 2021
Surgical Simulation and Training被引用数 6
ひとこと要約

本研究では、33件の腹腔鏡胆嚢切除手術動画からなるマルチセンターデータセット「HeiCholeベンチマーク」を用いて、手術プロセスとスキル分析のための機械学習アルゴリズムの評価が行われた。フェーズ認識では優れた性能(F1最高67.7%)を示したが、アクション認識とスキル評価は依然として困難であり、F1スコアは21.8–23.3%、平均誤差は0.78と高く、分野はまだ解決されていないものの、有望な可能性を秘めていることが示された。

ABSTRACT

PURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center dataset. In this work we investigated the generalizability of phase recognition algorithms in a multi-center setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 hours was created. Labels included annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 teams submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n=9 teams), for instrument presence detection between 38.5% and 63.8% (n=8 teams), but for action recognition only between 21.8% and 23.3% (n=5 teams). The average absolute error for skill assessment was 0.78 (n=1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but are not solved yet, as shown by our comparison of algorithms. This novel benchmark can be used for comparable evaluation and validation of future work.

研究の動機と目的

  • 単一センターのデータセットにとどまらない、手術フェーズ認識アルゴリズムの一般化性能を評価すること。
  • マルチセンターモデル環境下で、手術アクション認識、器具検出、スキル評価のための機械学習モデルを評価すること。
  • 手術AIシステムの比較的評価のための標準化されたベンチマークを確立すること。
  • 現在の手術プロセスおよびスキル分析アルゴリズムにおける性能ギャップを特定すること。

提案手法

  • 3つの外科センターから収集した33件の腹腔鏡胆嚢切除手術動画からなるマルチセンターデータセットを構築し、合計22時間の手術時間のデータを取得した。
  • アノテーションには、7つの手術フェーズ(250回の遷移)、5,514件の手術アクション、21種類の器具タイプにわたる6,980件の器具インスタンス、5次元にわたる495件のスキル評価が含まれた。
  • このデータセットは2019年エンドスコピックビジョンチャレンジのサブチャレンジ「手術プロセスとスキル分析」で使用され、12チームがアルゴリズムを提出した。
  • アルゴリズムの評価には、フェーズ、アクション、器具認識のF1スコアと、スキル評価の平均絶対誤差が用いられた。
  • 実際の手術環境における耐性と一般化性能を評価するため、参加チーム間での性能比較が行われた。
  • HeiCholeベンチマークは、今後の研究のための標準化された評価プラットフォームとして確立された。

実験結果

リサーチクエスチョン

  • RQ1手術フェーズ認識のための機械学習モデルは、複数の外科センター間でどれほど一般化できるか?
  • RQ2フェーズや器具認識と比較して、モデルの手術アクション認識性能はどの程度か?
  • RQ3機械学習モデルは、複数の次元にわたって手術スキルを正確に評価できるか?
  • RQ4現在の手術プロセスおよびスキル分析システムにおける主な制限要因は何か?
  • RQ5HeiCholeベンチマークは、手術AIアルゴリズムの公平かつ比較可能な評価をどのように可能にするか?

主な発見

  • 9チームのフェーズ認識ではF1スコアが23.9%から67.7%の間で変動し、性能にばらつきが見られ、改善の余地があることが示された。
  • 8チームの器具存在検出ではF1スコアが38.5%から63.8%の間で変動し、中程度の性能ではあるが一貫性に欠ける結果となった。
  • アクション認識の性能が最も低く、5チームのF1スコアは21.8%から23.3%の間で推移した。
  • スキル評価の平均絶対誤差は0.78と高く、定量的スキル評価の著しい不正確さが示された。
  • HeiCholeベンチマークは、現在のアルゴリズムが実際の手術現場への導入に耐えうるほど堅牢で信頼性がないことを明らかにした。
  • 本研究は、技術的潜在能力が著しく高いものの、手術プロセスとスキル分析が依然として未解決の課題であることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。