Skip to main content
QUICK REVIEW

[논문 리뷰] Disruptive Autoencoders: Leveraging Low-level features for 3D Medical Image Pre-training

Jeya Maria Jose Valanarasu, Yucheng Tang|arXiv (Cornell University)|2023. 07. 31.
Radiomics and Machine Learning in Medical Imaging인용 수 5
한 줄 요약

이 논문은 3D 의료 영상에 대한 새로운 자기지도 학습 사전 훈련 프레임워크인 교란 자동에코더(Disruptive Autoencoders, DAE)를 제안한다. 이는 채널 임베딩에 대해 마스킹을 수행하는 국소 마스킹과 소음, 다운샘플링과 같은 저수준 편향을 조합하여 저수준 특징 학습을 향상시킨다. 이 방법은 BTCV 다기관 분할 도전 대회 공개 테스트 리더보드에서 최고 성능을 기록하며 여러 3D 의료 영상 분할 벤치마크에서 최신 기술 성능(SOTA)을 달성한다.

ABSTRACT

Harnessing the power of pre-training on large-scale datasets like ImageNet forms a fundamental building block for the progress of representation learning-driven solutions in computer vision. Medical images are inherently different from natural images as they are acquired in the form of many modalities (CT, MR, PET, Ultrasound etc.) and contain granulated information like tissue, lesion, organs etc. These characteristics of medical images require special attention towards learning features representative of local context. In this work, we focus on designing an effective pre-training framework for 3D radiology images. First, we propose a new masking strategy called local masking where the masking is performed across channel embeddings instead of tokens to improve the learning of local feature representations. We combine this with classical low-level perturbations like adding noise and downsampling to further enable low-level representation learning. To this end, we introduce Disruptive Autoencoders, a pre-training framework that attempts to reconstruct the original image from disruptions created by a combination of local masking and low-level perturbations. Additionally, we also devise a cross-modal contrastive loss (CMCL) to accommodate the pre-training of multiple modalities in a single framework. We curate a large-scale dataset to enable pre-training of 3D medical radiology images (MRI and CT). The proposed pre-training framework is tested across multiple downstream tasks and achieves state-of-the-art performance. Notably, our proposed method tops the public test leaderboard of BTCV multi-organ segmentation challenge.

연구 동기 및 목표

  • 3D 의료 영상에서 임상적 세분화 작업(예: 기관 및 병변 분할)에 필수적인 강력한 저수준 표현을 학습하는 데 도전하는 것.
  • 기본 마스킹 자동에코더(Masked Autoencoders, MAEs)가 3D 방사선 영상에서 세밀한 해부학적 구조를 재구성하는 데 한계를 보이는 문제를 해결하는 것.
  • 여러 의료 영상 모odalities(MRI, CT 등)를 통합된 대비 학습 인식 방식으로 효과적으로 활용할 수 있는 사전 훈련 프레임워크를 설계하는 것.
  • 효과적인 사전 훈련을 가능하게 하기 위해 대규모 다중 모odal 3D 의료 영상 데이터셋을 구축하고 활용하는 것.
  • 더 강력한 저수준 특징 학습이 의료 영상 분할 작업의 최종 성능 향상에 기여함을 입증하는 것.

제안 방법

  • 공간 토큰이 아닌 채널 임베딩을 대상으로 마스킹하는 새로운 국소 마스킹 전략을 제안하여, 모델이 국소적이고 세밀한 특징을 학습하는 능력을 향상시킨다.
  • 토큰화 이전에 적용되는 교란 편향(추가 소음 및 다운샘플링)을 도입하여 저수준 표현 학습을 더욱 자극한다.
  • 이러한 교란 요소들을 트랜스포머 기반 인코더와 컨volutional 디코더를 조합하여 손상된 잠재 표현에서 원본 3D 영상을 재구성하는 방식을 구현한다.
  • 사전 훈련 중 다양한 영상 모달리티(MRI, CT 등) 간 특징을 정렬하기 위해 교차 모달 대비 손실(Cross-Modal Contrastive Loss, CMCL)을 도입하여 모달리티 인식 기반 표현 학습을 가능하게 한다.
  • 다양한解부학적 및 촬영 조건의 변형을 포함한 대규모로 구축된 3D MRI 및 CT 볼륨 데이터셋을 활용해 모델을 사전 훈련한다.
  • 자기지도 사전 훈련을 수행한 후 표준 벤치마크(예: BTCV 및 PROMISE12)에서의 표준 분할 작업에 대해 미세 조정(fine-tuning)을 수행한다.
Figure 1: Disruptive Autoencoders: Here, we disrupt the 3D medical image with a combination of low-level perturbations - noise and downsampling, followed by tokenization and local masking. These disrupted tokens are then passed through a transformer encoder and convolutional decoder to learn to reco
Figure 1: Disruptive Autoencoders: Here, we disrupt the 3D medical image with a combination of low-level perturbations - noise and downsampling, followed by tokenization and local masking. These disrupted tokens are then passed through a transformer encoder and convolutional decoder to learn to reco

실험 결과

연구 질문

  • RQ1목표적인 교란을 통해 저수준 특징 학습에 초점을 맞춘 사전 훈련 프레임워크가 3D 의료 영상 표현 학습에 향상 효과를 줄 수 있는가?
  • RQ2공간 토큰이 아닌 채널 임베딩을 마스킹하는 국소 마스킹 방식이 기존 토큰 마스킹보다 더 나은 재구성 및 최종 성능을 내는가?
  • RQ3저수준 편향(소음, 다운샘플링)과 국소 마스킹을 병합한 전략이 개별적인 교란 요소보다 사전 훈련 효과가 뛰어나게 되는가?
  • RQ4교차 모달 대비 학습이 3D 의료 영상에서 다중 모달 사전 훈련에 얼마나 기여하는가?
  • RQ5더 강력한 저수준 특징 학습이 더 낮은 데이터 환경에서도 일반화 능력을 향상시키는가?

주요 결과

  • 교란 자동에코더는 BTCV 다기관 분할 도전 대회 공개 테스트 리더보드에서 이전 방법들(MAE 및 Swin-UNETR 포함)을 모두 능가하며 최신 기술 성능(SOTA)을 달성했다.
  • BTCV 데이터셋에서 DAE는 딱시 스코어 0.8472를 기록하여, Swin-UNETR의 세 가지 프레텍스트 작업 조합을 초월했다.
  • PROMISE12 데이터셋에서 DAE는 모든 데이터 환경에서 랜덤 초기화 및 MAE를 모두 능가했으며, 20% 및 50% 훈련 데이터에서 뚜렷한 성능 향상을 보여, 높은 데이터 효율성을 입증했다.
  • 절단 분석 결과, 국소 마스킹과 저수준 편향의 조합이 가장 뛰어난 성능을 내며, 국소 마스킹만으로도 소음 및 다운샘플링을 개별적으로 적용한 것보다 성능이 뛰어나다는 것이 확인되었다.
  • CKA 분석 결과, 사전 훈련된 모델에서 미세 조정된 모델로 이르기까지 초기 네트워크 레이어가 더 많은 저수준 특징을 유지함을 확인하여, 저수준 표현이 더 안정적이고 최종 성능에 핵심적임을 입증했다.
  • 교차 모달 대비 손실(CMCL)의 포함으로 인해 미세 조정 성능이 향상되었으며, 이는 모달리티 인식 기반 사전 훈련이 다양한 영상 모달리티 간 일반화 능력을 향상시킨다는 것을 시사한다.
Figure 2: Comparison of reconstruction quality. It can be observed that masked image modelling produces a coarse reconstruction for radiology without local context while the proposed disruptions in this work obtain sharper reconstructions recovering meaningful fine details.
Figure 2: Comparison of reconstruction quality. It can be observed that masked image modelling produces a coarse reconstruction for radiology without local context while the proposed disruptions in this work obtain sharper reconstructions recovering meaningful fine details.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.