Skip to main content
QUICK REVIEW

[논문 리뷰] On Measuring and Controlling the Spectral Bias of the Deep Image Prior

Zenglin Shi, Pascal Mettes|arXiv (Cornell University)|2021. 07. 02.
Photoacoustic and Ultrasonic Imaging인용 수 5
한 줄 요약

이 논문은 딥 이미지 프리독의 스펙트럴 바이어스(저주파 성분이 고주파 성분보다 빠르고 잘 학습됨)를 측정하고 제어하기 위해 주파수 대역 대응 지표를 도입한다. 리프시츠 제어 컨볼루션과 가우시안 제어 업샘플링 레이어를 통합함으로써 최적화 중 성능 저하를 방지하고 자동 정지 기능을 가능하게 하여 오рак루 정지 기준이 필요 없이 이미지 노이즈 제거, 디블로킹, 인painting, 슈퍼레졸루션, 디테일 강화 등에서 최신 기술 수준의 성능을 달성한다.

ABSTRACT

The deep image prior showed that a randomly initialized network with a suitable architecture can be trained to solve inverse imaging problems by simply optimizing it's parameters to reconstruct a single degraded image. However, it suffers from two practical limitations. First, it remains unclear how to control the prior beyond the choice of the network architecture. Second, training requires an oracle stopping criterion as during the optimization the performance degrades after reaching an optimum value. To address these challenges we introduce a frequency-band correspondence measure to characterize the spectral bias of the deep image prior, where low-frequency image signals are learned faster and better than high-frequency counterparts. Based on our observations, we propose techniques to prevent the eventual performance degradation and accelerate convergence. We introduce a Lipschitz-controlled convolution layer and a Gaussian-controlled upsampling layer as plug-in replacements for layers used in the deep architectures. The experiments show that with these changes the performance does not degrade during optimization, relieving us from the need for an oracle stopping criterion. We further outline a stopping criterion to avoid superfluous computation. Finally, we show that our approach obtains favorable results compared to current approaches across various denoising, deblocking, inpainting, super-resolution and detail enhancement tasks. Code is available at \url{https://github.com/shizenglin/Measure-and-Control-Spectral-Bias}.

연구 동기 및 목표

  • 딥 이미지 프리독에서 저주파 이미지 성분이 고주파 성분보다 더 쉽게 학습됨을 이해하고 정량적으로 측정하는 것.
  • 네트워크 아키텍처 선택 외의 제어 수단이 부족한 문제를 해결하는 것.
  • 고주파 노이즈에 대한 과적합을 방지함으로써 최적화 중 성능 저하를 방지하는 것.
  • 불필요한 계산을 피하기 위해 신뢰할 수 있고 자동으로 정지되는 기준을 개발하는 것.
  • 다양한 역영상 처리 작업에서 수렴 속도와 복원 품질을 향상시키는 것.

제안 방법

  • 최적화 중 다양한 주파수 대역의 학습 속도를 비교함으로써 스펙트럴 바이어스를 정량적으로 측정하기 위해 주파수 대역 대응 지표를 도입한다.
  • 기울기 흐름을 제약하여 고주파 적합을 제한하고 훈련을 안정화시키는 리프시츠 제어 컨볼루션 레이어를 제안한다.
  • 조절 가능한 스무딩을 허용하여 주파수 응답과 수렴 속도를 균형 잡는 가우시안 제어 업샘플링 레이어를 도입한다.
  • 추가로 최적화 안정성과 일반화 성능 향상을 위해 리프시츠 제약을 적용한 수정된 배치 정규화 레이어를 사용한다.
  • 정점 성능에 도달할 때 최적화를 정지시키기 위해 스펙트럴 바이어스 모니터링 기반의 자동 정지 기준을 개발한다.
  • 이미지 강화 작업에서 부드러움을 제어하기 위해 조정 가능한 하이퍼파ram터(예: 식 (5)의 λ)를 사용하여 고정된 5,000회 반복을 수행한다.
Figure 1: Frequency-band correspondence metric. The left image shows an example of correspondence map $H$ , which is computed according to Eq. ( 1 ). We divide the correspondence map into $N$ subgroups corresponding to $N$ non-overlapping frequency bands. Since the correspondence map is symmetrical
Figure 1: Frequency-band correspondence metric. The left image shows an example of correspondence map $H$ , which is computed according to Eq. ( 1 ). We divide the correspondence map into $N$ subgroups corresponding to $N$ non-overlapping frequency bands. Since the correspondence map is symmetrical

실험 결과

연구 질문

  • RQ1딥 이미지 프리독 최적화 과정에서 스펙트럴 바이어스는 어떻게 나타나며, 이를 정량적으로 측정할 수 있는가?
  • RQ2아키텍처 수정을 통해 스펙트럴 바이어스를 제어할 수 있는가? 성능 저하를 방지할 수 있는가?
  • RQ3오라클 정지 기준에 의존하지 않고도 고주파 노이즈에 대한 과적합으로부터 딥 이미지 프리독을 강건하게 만들 수 있는가?
  • RQ4스펙트럴 행동 기반으로 자동 정지 기준을 설계할 수 있는가? 이를 통해 불필요한 계산을 줄일 수 있는가?
  • RQ5스펙트럴 바이어스를 제어하면 다양한 역영상 처리 작업에서 성능 향상이 이루어지는가?

주요 결과

  • 이 방법은 이미지 노이즈 제거 작업에서 Set14 데이터셋에서 PSNR 32.76을 달성하여 Ulyanov 등(2020)과 LapSRN을 포함한 이전 방법들을 능가한다.
  • 8배 확대 슈퍼레졸루션 작업에서 PSNR 30.45를 기록하여 이전 최고 성능인 LapSRN의 28.54를 초월한다.
  • 최적화 중 성능 저하가 완전히 제거되어 오라클 정지 기준 없이도 정점 성능을 유지한다.
  • 자동 정지 기준은 불필요한 계산을 효과적으로 줄이며 높은 품질의 결과를 유지한다.
  • 이미지 강화 작업에서는 λ 하이퍼파ram터를 조정함으로써 샤프닝 강도를 정밀하게 제어할 수 있어 일관된 디테일 강화가 가능하다.
  • 고주파 노이즈 수준이 높을 경우(σ=100), 낮은 주파수에서 노이즈와 자연 이미지의 전력 스펙트럼이 겹쳐져 주파수 분리가 어려워지기 때문에 실패한다.
Figure 2: Network architectures used in the experiments of Section 3 . The Encoder-Decoder is the same as the one used in Ulyanov et al. ( 2020 ) . Specifically, the encoder contains five convolution blocks. Each block contains two convolution layers with the kernel size of $3\times 3$ and the chann
Figure 2: Network architectures used in the experiments of Section 3 . The Encoder-Decoder is the same as the one used in Ulyanov et al. ( 2020 ) . Specifically, the encoder contains five convolution blocks. Each block contains two convolution layers with the kernel size of $3\times 3$ and the chann

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.