Skip to main content
QUICK REVIEW

[논문 리뷰] Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks

Gaurav Yengera, Didier Mutter|arXiv (Cornell University)|2018. 05. 22.
Reservoir Engineering and Simulation Methods참고 문헌 37인용 수 51
한 줄 요약

수술 단계 인식을 위한 반지도학습 접근법을 제시하며, Remaining Surgery Duration(RSD)에서 자체 감독 사전 학습으로 CNN-LSTM 네트워크를 학습시키고, RSD 사전 학습을 통한 엔드투엔드 학습이 주석 데이터 필요를 줄이면서 성능을 유지하거나 향상시킬 수 있음을 보인다.

ABSTRACT

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve safety within the OR. State-of-the-art surgical phase recognition algorithms using laparoscopic videos are based on fully supervised training. This limits their potential for widespread application, since creation of manual annotations is an expensive process considering the numerous types of existing surgeries and the vast amount of laparoscopic videos available. In this work, we propose a new self-supervised pre-training approach based on the prediction of remaining surgery duration (RSD) from laparoscopic videos. The RSD prediction task is used to pre-train a convolutional neural network (CNN) and long short-term memory (LSTM) network in an end-to-end manner. Our proposed approach utilizes all available data and reduces the reliance on annotated data, thereby facilitating the scaling up of surgical phase recognition algorithms to different kinds of surgeries. Additionally, we present EndoN2N, an end-to-end trained CNN-LSTM model for surgical phase recognition and evaluate the performance of our approach on a dataset of 120 Cholecystectomy laparoscopic videos (Cholec120). This work also presents the first systematic study of self-supervised pre-training approaches to understand the amount of annotations required for surgical phase recognition. Interestingly, the proposed RSD pre-training approach leads to performance improvement even when all the training data is manually annotated and outperforms the single pre-training approach for surgical phase recognition presently published in the literature. It is also observed that end-to-end training of CNN-LSTM networks boosts surgical phase recognition performance.

연구 동기 및 목표

  • 복강경 비디오에서 수술 단계 인식에 대한 수동으로 주석된 데이터 의존도를 줄인다.
  • 대규모의 비라벨링 비디오를 활용하기 위해 자체 감독 사전 학습을 활용한다.
  • 수술 단계 및 절차 전반에서 일반화를 개선하기 위해 엔드투엔드 CNN-LSTM 학습을 촉진한다.

제안 방법

  • EndoN2N 도입: 전체 비디오 시퀀스에 대해 학습되며 시간에 따른 근사적 역전파를 사용하는 엔드투엔드 CNN-LSTM 모델.
  • 수술 비디오로부터 Remaining Surgery Duration(RSD)를 예측하여 자체 감독 사전 학습을 수행하고 CNN-LSTM 모델을 초기화한다.
  • 동일 아키텍처에서 EndoN2N과 두 단계 CNN-그 다음 LSTM 학습 방식인 EndoLSTM을 비교한다.
  • 수술 단계 인식을 위한 엔드투엔드와 두 단계 학습 간의 공정한 비교를 제공한다.
  • RSD 사전 학습 및 단계 인식 중에 남은 수술 시간(RSD)과 진행 상황을 다중 작업 신호로 통합한다.
  • 메모리 제약으로 인해 긴 비디오 시퀀스를 훈련을 위해 부분 시퀀스로 분할하고 경계 상태 전파를 수행한다.

실험 결과

연구 질문

  • RQ1주석 데이터의 양이 달라질 때 자체 감독 RSD 사전 학습이 수술 단계 인식 성능에 어떤 영향을 미치는가?
  • RQ2전체 비디오 시퀀스에서 엔드투엔드 CNN-LSTM 학습이 두 단계의 EndoLSTM 접근보다 수술 단계 인식에서 더 우수한가?
  • RQ3RSD 사전 학습이 서로 다른 수술에도 일반화되고 확장 가능한 반지도학습을 지원할 수 있는가?

주요 결과

  • RSD 사전 학습은 주석 비디오를 약 20% 정도 줄이면서도 유사하거나 약간 더 나은 수술 단계 인식을 가능하게 한다.
  • 주석 비디오 약 50% 더 적은 경우에도 성능 차이는 대략 5% 이내이다.
  • EndoN2N으로의 엔드투엔드 CNN-LSTM 학습은 두 단계 학습(EndoLSTM)에 비해 성능과 일반화를 개선한다.
  • 모든 학습 데이터가 완전히 주석되어 있어도 RSD 사전 학습은 성능 향상을 제공하며, 이전에 발표된 단일 사전 학습 접근법보다 우수하다.
  • 제안된 근사치를 사용하여 긴 비디오 시퀀스에서 CNN-LSTM을 엔드투엔드로 학습하는 것이 가능하며 더 나은 결과를 산출한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.