Skip to main content
QUICK REVIEW

[논문 리뷰] AT-DDPM: Restoring Faces degraded by Atmospheric Turbulence using Denoising Diffusion Probabilistic Models

Nithin Gopalakrishnan Nair, Kangfu Mei|arXiv (Cornell University)|2022. 08. 24.
Advanced Image Processing Techniques인용 수 9
한 줄 요약

이 논문은 대기 난류로 인해 손상된 얼굴 영상을 복원하기 위한 첫 번째 디노이징 확산 확률 모델(DDPM) 기반 방법인 AT-DDPM을 제안한다. 깨끗한 얼굴에 사전 훈련된 DDPM을 활용하고 점진적 훈련을 통해 지식 distillation을 적용함으로써, 새로운 효율적인 샘플링 전략을 통해 왜곡된 이미지에서 시작하여 현실적이고 고해상도의 얼굴을 재구성하는 모델이 되었으며, 합성 및 실제 데이터셋에서 최고 성능을 달성하였다.

ABSTRACT

Although many long-range imaging systems are designed to support extended vision applications, a natural obstacle to their operation is degradation due to atmospheric turbulence. Atmospheric turbulence causes significant degradation to image quality by introducing blur and geometric distortion. In recent years, various deep learning-based single image atmospheric turbulence mitigation methods, including CNN-based and GAN inversion-based, have been proposed in the literature which attempt to remove the distortion in the image. However, some of these methods are difficult to train and often fail to reconstruct facial features and produce unrealistic results especially in the case of high turbulence. Denoising Diffusion Probabilistic Models (DDPMs) have recently gained some traction because of their stable training process and their ability to generate high quality images. In this paper, we propose the first DDPM-based solution for the problem of atmospheric turbulence mitigation. We also propose a fast sampling technique for reducing the inference times for conditional DDPMs. Extensive experiments are conducted on synthetic and real-world data to show the significance of our model. To facilitate further research, all codes and pretrained models are publically available at http://github.com/Nithin-GK/AT-DDPM

연구 동기 및 목표

  • 대기 난류로 인해 블러 및 기하학적 왜곡이 발생하는 얼굴 영상 복원 문제를 해결하기 위해.
  • 기존의 CNN 및 GAN 기반 방법의 한계인 불안정한 훈련, 모드 붕괴, 실제 데이터에서의 낮은 일반화 능력을 극복하기 위해.
  • 디노이징 확산 확률 모델(DDPM)이 이미지 복원을 위한 안정적이고 고품질의 생성 사전으로서의 잠재력을 탐색하기 위해.
  • 순수한 노이즈가 아닌 왜곡된 이미지에서 시작하는 새로운 샘플링 전략을 통해 DDPM의 일반적인 느린 추론 시간을 줄이기 위해.
  • 지속적 학습 기반 지식 distillation 프레임워크를 통해 고난류 조건에서도 재구성 정확도와 신원 유지 능력을 향상시키기 위해.

제안 방법

  • 대규모 깨끗한 얼굴 데이터셋에서 사전 훈련된 DDPM을 지식 distillation을 통해 미세조정하여 단일 영상의 얼굴 초해상도 복원을 수행한다.
  • 초해상도 모델은 추가로 지식 distillation을 통해 왜곡된 입력에 대한 깨끗한 얼굴의 조건부 분포를 학습하도록 적응시킨다.
  • 점진적 훈련(PT) 프레임워크를 도입하여 모델을 복원 작업에 점차적으로 적응시킴으로써 훈련의 안정성과 재구성 품질을 향상시킨다.
  • 역확산 과정을 순수한 가우시안 노이즈가 아닌 왜곡된 입력 영상에서 시작하는 효율적인 샘플링 전략을 제안하여 신원 일관성을 향상시킨다.
  • 강한 난류 조건에서도 얼굴의 구조적 특징이 입력과 유사하게 유지되도록 특징 일관성을 강제함으로써 재구성 정확도를 향상시킨다.
  • 학습된 현실적 얼굴의 다양체를 활용하여 입력이 심하게 손상된 경우에도 사진 수준의 정교한 출력을 생성한다.
Figure 2 : An overview of the proposed approach, during the training process, we perform knowledge distillation to transfer class prior information from a network trained for image super-resolution on a large dataset to the network for removing turbulence degradation. During inference, rather than s
Figure 2 : An overview of the proposed approach, during the training process, we perform knowledge distillation to transfer class prior information from a network trained for image super-resolution on a large dataset to the network for removing turbulence degradation. During inference, rather than s

실험 결과

연구 질문

  • RQ1DDPM는 단일 영상의 대기 난류 완화를 위한 생성 사전으로 효과적으로 적응시킬 수 있는가?
  • RQ2점진적 지식 distillation은 난류 입력에 대한 DDPM 기반 이미지 복원의 품질과 안정성에 어떻게 기여하는가?
  • RQ3순수한 노이즈가 아닌 왜곡된 이미지에서 역확산 과정을 시작하면 신원 유지 능력 향상과 추론 시간 단축에 기여하는가?
  • RQ4제안된 방법은 GAN 기반 및 CNN 기반 접근법과 비교해 실제 데이터에서 시각적 품질, 정확도, 내성에 대해 어떻게 우월한가?
  • RQ5효율적인 샘플링 전략은 추론 속도와 출력 일관성에 어떤 영향을 미치는가?

주요 결과

  • 제안된 AT-DDPM 모델은 합성(LRFID) 및 실제 데이터셋(BRIAR) 모두에서 최고 성능을 달성하였으며, 시각적 품질과 구조적 정확도에서 기존 방법을 능가하였다.
  • LRFID 데이터셋에서 점진적 훈련을 적용한 결과, 상위 1위 얼굴 인식 정확도가 54.8%에 도달하였으며, 이는 점진적 훈련 없이 48.7%를 기록한 것에 비해 뚜렷한 향상이다.
  • 점진적 훈련 없이 6.985였던 NIQE(비기반 이미지 품질 측정 지표)는 6.824로 감소하여 인지적 품질 향상을 나타낸다.
  • LPIPS 점수는 0.5358에서 0.5255로 감소하여 재구성된 영상와 진짜 영상 간의 인지적 유사도가 향상됨을 보여준다.
  • 효율적인 샘플링 전략은 추론 시간을 단축시키면서도 고품질 출력을 유지함을 오차 막대 분석을 통해 검증하였다.
  • 정성적 결과에서는 GFPGAN, LTTGAN, ATNet과 비교해 AT-DDPM가 더 나은 얼굴 신원 유지 능력을 보이며 색조 이동과 잡음 생성을 방지함을 확인하였다.
(a) Distorted
(a) Distorted

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.