[논문 리뷰] Blind Deblurring Using GANs
이 논문은 주의 모듈, 잔차 연결, 다중 손실 훈련을 통해 전역적 인식을 향상시켜 블라인드 이미지 디블러링을 위한 GAN 기반 접근법을 제안한다. 이 방법은 GoPro 데이터셋에서 성능을 향상시키며, DeblurGAN에서 PSNR 28.70과 SSIM 0.958를 달성했고, 자기주의 주의와 인지적 손실을 사용함으로써 추가적인 성능 향상을 이룬다.
Deblurring is the task of restoring a blurred image to a sharp one, retrieving the information lost due to the blur. In blind deblurring we have no information regarding the blur kernel. As deblurring can be considered as an image to image translation task, deep learning based solutions, including the ones which use GAN (Generative Adversarial Network), have been proven effective for deblurring. Most of them have an encoder-decoder structure. Our objective is to try different GAN structures and improve its performance through various modifications to the existing structure for supervised deblurring. In supervised deblurring we have pairs of blurred and their corresponding sharp images, while in the unsupervised case we have a set of blurred and sharp images but their is no correspondence between them. Modifications to the structures is done to improve the global perception of the model. As blur is non-uniform in nature, for deblurring we require global information of the entire image, whereas convolution used in CNN is able to provide only local perception. Deep models can be used to improve global perception but due to large number of parameters it becomes difficult for it to converge and inference time increases, to solve this we propose the use of attention module (non-local block) which was previously used in language translation and other image to image translation tasks in deblurring. Use of residual connection also improves the performance of deblurring as features from the lower layers are added to the upper layers of the model. It has been found that classical losses like L1, L2, and perceptual loss also help in training of GANs when added together with adversarial loss. We also concatenate edge information of the image to observe its effects on deblurring. We also use feedback modules to retain long term dependencies
연구 동기 및 목표
- 비균일한 블러를 처리하기 위해 전체 이미지에 걸친 맥락적 이해가 필요한 블라인드 이미지 디블러링 모델에서 전역적 인식을 향상시키기.
- CNN의 국소적 수신장의 한계를 극복하기 위해 주의 메커니즘을 통합하여 장거리 종속성을 모델링하기.
- 대비 손실과 L1, L2, 인지적 손실을 조합하여 훈련 안정성과 성능을 향상시키기.
- 잔차 연결, 에지 정보, 피드백 모듈이 디블러링 품질 향상에 미치는 영향을 평가하기.
- 더 나은 선명도와 구조적 유사도를 확보하기 위해 아키텍처 수정과 손실 함수를 통해 기존 GAN 아키텍처(Pix2Pix, DeblurGAN, RiR)를 최적화하기.
제안 방법
- 모든 세 개의 인코더/디코더 블록 이후에 비국소 자기주의 주의 모듈을 통합하여 장거리 종속성을 포착하고 전역적 맥락 모델링을 향상시키기.
- 채널 기반 주의를 적용하여 특징 맵을 재조정하고, 주목할 만한 채널에 집중함으로써 표현 학습을 향상시키기.
- 전역 잔차 연결을 사용하여 깊은 레이어를 통해 저수준 특징을 유지함으로써 기울기 흐름과 특징 재사용을 지원하기.
- 대비 손실과 VGG 특징 기반 인지적 손실, L1 및 L2 손실을 조합하여 훈련 안정성과 시각적 품질을 향상시키기.
- GAN 훈련의 안정성을 높이고 모드 붕괴를 방지하기 위해 디스criminator에 스펙트럼 정규화를 구현하기.
- 다중 스텝을 거쳐 디블러딩 출력을 반복적으로 개선하는 피드백 모듈을 통합하여 장기 종속성을 유지하기.
실험 결과
연구 질문
- RQ1주의 메커니즘이 블라인드 이미지 디블러링에서 GAN의 전역적 인식에 상당한 영향을 미칠 수 있는가?
- RQ2잔차 연결과 특징 재사용은 깊은 인코더-디코더 아키텍처에서 디블러딩 성능에 어떤 영향을 미치는가?
- RQ3대비 손실, 인지적 손실, L1/L2 손실을 병합하면 수렴성과 더 선명한 출력을 향상시키는가?
- RQ4에지 정보와 피드백 모듈은 GAN 기반 모델에서 디블러딩 품질에 어떤 영향을 미치는가?
- RQ5심화된 네트워크, 스킵 연결 등의 아키텍처 선택은 블라인드 디블러링에서 PSNR와 SSIM에 어떤 영향을 미치는가?
주요 결과
- DeblurGAN 모델은 GoPro 데이터셋에서 PSNR 28.70과 SSIM 0.958를 달성하여 기준 모델을 초월했다.
- Pix2Pix에 자기주의 주의와 인지적 손실을 추가함으로써 PSNR는 25.41에서 27.29로, SSIM은 0.810에서 0.858로 향상되었다.
- 주의와 인지적 손실을 통합한 Residual-in-Residual(RiR) 모델은 고해상도 1280×768 이미지에서 PSNR 23.46을 기록했다.
- 피드백 모듈은 성능을 떨어뜨렸으며, PSNR는 27.20으로, SSIM은 0.827로 하락하여 이 설정에서 유의미한 이점이 없음을 시사했다.
- 에지 정보 통합은 성능을 악화시켜 PSNR는 25.27로, SSIM은 0.773으로 감소하여 이 작업에 효과적이지 않음을 시사했다.
- 스펙트럼 정규화와 다중 손실 훈련은 모든 아키텍처에서 훈련 안정성과 최종 모델 성능을 크게 향상시켰다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.