[論文レビュー] Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks
Remaining Surgery Duration (RSD) を用いた自己監督型前訓練による CNN-LSTM ネットワークを用いた外科フェーズ認識の半教師付きアプローチを提案する;RSD によるエンドツーエンド訓練が注釈データ量を減らす一方で性能を維持または向上させることを示す。
Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve safety within the OR. State-of-the-art surgical phase recognition algorithms using laparoscopic videos are based on fully supervised training. This limits their potential for widespread application, since creation of manual annotations is an expensive process considering the numerous types of existing surgeries and the vast amount of laparoscopic videos available. In this work, we propose a new self-supervised pre-training approach based on the prediction of remaining surgery duration (RSD) from laparoscopic videos. The RSD prediction task is used to pre-train a convolutional neural network (CNN) and long short-term memory (LSTM) network in an end-to-end manner. Our proposed approach utilizes all available data and reduces the reliance on annotated data, thereby facilitating the scaling up of surgical phase recognition algorithms to different kinds of surgeries. Additionally, we present EndoN2N, an end-to-end trained CNN-LSTM model for surgical phase recognition and evaluate the performance of our approach on a dataset of 120 Cholecystectomy laparoscopic videos (Cholec120). This work also presents the first systematic study of self-supervised pre-training approaches to understand the amount of annotations required for surgical phase recognition. Interestingly, the proposed RSD pre-training approach leads to performance improvement even when all the training data is manually annotated and outperforms the single pre-training approach for surgical phase recognition presently published in the literature. It is also observed that end-to-end training of CNN-LSTM networks boosts surgical phase recognition performance.
研究の動機と目的
- 腹腔鏡ビデオにおける手術フェーズ認識のための手動で注釈されたデータへの依存を減らす。
- 自己監督型前学習を活用して大規模なラベルなしビデオを利用する。
- 外科フェーズと手技間での汎化性能を向上させるために、エンドツーエンドの CNN-LSTM 訓練を推進する。
提案手法
- EndoN2N を導入する。これは完全なビデオ列に対して近似的な time- backpropagation through time を用いたエンドツーエンドの CNN-LSTM モデルである。
- 腹腔鏡ビデオから Remaining Surgery Duration (RSD) を予測することで自己監督型事前学習を行い、CNN-LSTM モデルを初期化する。
- 同一アーキテクチャで EndoN2N を、先に CNN を学習しその後 LSTM を訓練する 2 段階の EndoLSTM アプローチと比較する。
- 外科フェーズ認識におけるエンドツーエンドと2 段階トレーニングの apples-to-apples な比較を提供する。
- RSD 前訓練時およびフェーズ認識時に、残り手術時間と進捗を多タスク信号として組み込む。
- メモリ制約のため、長いビデオ列を訓練用にサブシーケンスに分割し、境界状態伝播を用いる。
実験結果
リサーチクエスチョン
- RQ1自己監督型 RSD 前訓練が、異なる量の注釈データに対して外科フェーズ認識の性能にどのように影響するか?
- RQ2完全なビデオ列に対して、エンドツーエンドの CNN-LSTM 訓練は2段階の EndoLSTM アプローチを上回るか?
- RQ3RSD 前訓練は異なる手術間で一般化し、スケーラブルな半教師付き学習をサポートできるか?
主な発見
- RSD 前訓練により、約20% fewer の注釈付きビデオで同程度またはわずかに優れた外科フェーズ認識が可能になる。
- 注釈付きビデオを約50%減らすと、性能差は約5%程度に収まる。
- CNN-LSTM ネットワークのエンドツーエンド訓練(EndoN2N)は、2段階訓練(EndoLSTM)と比較して性能と汎化を向上させる。
- すべての訓練データが完全に注釈済みの場合でも、RSD 前訓練は性能を向上させ、外科フェーズ認識のための従来の単一前訓練法を上回る。
- 提案された近似法で長いビデオ列上で CNN-LSTM をエンドツーエンド訓練することは実現可能で、改善された結果をもたらす。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。