Skip to main content
QUICK REVIEW

[논문 리뷰] Multi-Scale Structural-aware Exposure Correction for Endoscopic Imaging

Axel Garcia-Vega, Ricardo Espinosa|arXiv (Cornell University)|2022. 10. 26.
Image Enhancement Techniques인용 수 4
한 줄 요약

이 논문은 LMSPEC 손실 함수에 구조적 유사성(SSIM)-aware 항을 추가하여 미세한 질감과 색상 유지 능력을 향상시킨 다중 스케일, 구조 인식형 딥러닝 방법을 제안한다. 이는 내 endoscopic 영상에서 실시간 曝光 보정을 위한 것이다. Endo4IE 데이터셋에서 과노출 및 과소노출 영상에 대해 각각 4.40% 및 4.21%의 상대적 SSIM 향상을 달성하였으며, 동시에 저지연 추론 시간을 유지한다.

ABSTRACT

Endoscopy is the most widely used imaging technique for the diagnosis of cancerous lesions in hollow organs. However, endoscopic images are often affected by illumination artefacts: image parts may be over- or underexposed according to the light source pose and the tissue orientation. These artifacts have a strong negative impact on the performance of computer vision or AI-based diagnosis tools. Although endoscopic image enhancement methods are greatly required, little effort has been devoted to over- and under-exposition enhancement in real-time. This contribution presents an extension to the objective function of LMSPEC, a method originally introduced to enhance images from natural scenes. It is used here for the exposure correction in endoscopic imaging and the preservation of structural information. To the best of our knowledge, this contribution is the first one that addresses the enhancement of endoscopic images using deep learning (DL) methods. Tested on the Endo4IE dataset, the proposed implementation has yielded a significant improvement over LMSPEC reaching a SSIM increase of 4.40% and 4.21% for over- and underexposed images, respectively.

연구 동기 및 목표

  • 인공지능 기반 진단 도구의 성능을 저하시키는 내시경 영상에서의 조명 아티팩트—과노출 및 과소노출—문제를 해결하기 위해.
  • 내시경 장면의 구조적 및 질감 세부 정보를 유지하면서 노출을 보정하는 실시간, 딥러닝 기반 영상 강화 방법을 개발하기 위해.
  • 기존 방법에서 유도된 질감 및 색상 아티팩트를 완화하기 위해 LMSPEC 프레임워크를 구조적 유사성 인식 손실 함수로 확장하기 위해.
  • Endo4IE 데이터셋에서 과노출 및 과소노출 영역의 정량적 지표와 정성적 충실도에 중점을 두고 제안된 방법을 평가하기 위해.
  • 컴퓨터 보조 내시경 시스템에 임상적 도입을 위한 강건하고 일반화 능력 있는 강화를 가능하게 하기 위해.

제안 방법

  • 노출 보정 과정에서 미세한 질감과 공간적 세부 정보의 유지에 우선순위를 두기 위해, LMSPEC 손실 함수에 구조적 유사성(SSIM) 항을 통합함으로써 이를 확장한다.
  • 두 단계 훈련 전략을 적용: 먼저 128×128 패치에서 훈련한 후, 전이된 가중치를 사용하여 256×256 패치에서 파인튜닝함으로써 특징 학습을 향상시킴.
  • 두 번째 단계에서 지정된 디스크림ิน레이터 시작 에포크(DSE)에 디스크림ิน레이터를 활성화하여 GAN 기반 훈련의 안정성과 현실감을 향상시킴.
  • 과소노출(UE), 과노출(OE), 병합(C) 데이터 서브셋 각각에 대해 별도의 모델을 훈련하고, 각 서브셋에 맞게 초모수를 튜닝하여 성능을 최적화함.
  • 최종 모델은 최적화된 학습률, 배치 크기, 훈련 에포크를 사용한 파인튜닝 설정을 활용하여 다양한 노출 유형 간의 일반화 능력을 향상시킴.
  • U-Net 기반 생성자와 적대적 훈련을 활용하며, 다중 스케일 SSIM 손실을 통해 인지적 품질을 유지함.
Figure 1 : Strong illumination change example in almost consecutive frames of a colonoscopic image sequence. (a) This image was acquired in appropriate lighting conditions. (b) Few frames later, the image is overexposed in its lower left region and underexposed in the remaining frame part.
Figure 1 : Strong illumination change example in almost consecutive frames of a colonoscopic image sequence. (a) This image was acquired in appropriate lighting conditions. (b) Few frames later, the image is overexposed in its lower left region and underexposed in the remaining frame part.

실험 결과

연구 질문

  • RQ1표준 LMSPEC에 비해 구조적 유사성 인식 손실 함수는 내시경 영상의 노출 보정에서 질감 및 세부 정보 유지에 있어 향상된 성능을 보일 수 있는가?
  • RQ2제안된 방법은 과노출 및 과소노출 내시경 영상에서 SSIM 및 PSNR 측면에서 LMSPEC보다 뛰어난 성능을 달성하는가?
  • RQ3임상 내시경에서 영상 품질 향상과 함께 실시간 추론 속도(예: 8 FPS 이상)를 유지할 수 있는가?
  • RQ4다양한 패치 크기에서의 다중 스케일 훈련은 모델의 다양한 조명 아티팩트에 대한 일반화 능력에 어떤 영향을 미치는가?
  • RQ5실제 내시경 영상 시퀀스에서 기준 LMSPEC에 비해 이 방법은 색상 및 질감 아티팩트를 얼마나 줄일 수 있는가?

주요 결과

  • 제안된 방법은 Endo4IE 데이터셋에서 과노출 영상에 대해 기준 LMSPEC 모델 대비 4.40%의 상대적 SSIM 향상을 달성하였다.
  • 과소노출 영상의 경우, LMSPEC 대비 4.21%의 상대적 SSIM 향상을 보이며, 미세한 구조적 세부 정보의 유지 능력 향상을 시사한다.
  • 최적화된 초모수를 적용한 병합 데이터(С)로 훈련된 모델이 모든 지표에서 LMSPEC 및 기준 모델을 초월하는 가장 높은 종합 성능을 달성하였다.
  • 정성적 결과는 제안된 방법이 LMSPEC에 비해 더 현실적이고 아티팩트가 없는 영상을 생성함을 보여주며, 특히 질감이 뚜렷한 영역에서 유의미한 개선을 보였다.
  • 다소의 개선에도 불구하고, 일부 강화된 영상에서 약간의 색조 이탈이 관찰되어 추가적인 색상 유지 손실 함수가 필요할 것으로 보인다.
  • 모델은 8 FPS의 추론 처리량을 달성하여 중간 수준의 실시간 성능을 보였으며, 임상적 도입을 위한 추가 최적화가 필요하다.
Figure 2 : DL-mode. On the left: Laplacian pyramid decomposition over patches I’ with exposure artefacts and Gaussian pyramid decomposition over ground truth patches T . On the right: $\pazocal{L}_{pyr}$ is computed with the up-sampled output from sub-networks 1 , 2 and 3 , whereas $\pazocal{L}_{rec
Figure 2 : DL-mode. On the left: Laplacian pyramid decomposition over patches I’ with exposure artefacts and Gaussian pyramid decomposition over ground truth patches T . On the right: $\pazocal{L}_{pyr}$ is computed with the up-sampled output from sub-networks 1 , 2 and 3 , whereas $\pazocal{L}_{rec

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.