Skip to main content
QUICK REVIEW

[논문 리뷰] Image Inpainting Guided by Coherence Priors of Semantics and Textures

Liang Liao, Jing Xiao|arXiv (Cornell University)|2020. 12. 15.
Generative Adversarial Networks and Image Synthesis참고 문헌 43인용 수 8
한 줄 요약

이 논문은 복잡한 시나리오에서 구멍 메우기에 있어 의미와 텍스처 간의 일관성 사전 지식을 활용하여 공동 이미지 보정 및 의미 분할 프레임워크를 제안한다. 의미 기반 주의 전파(SWAP) 모듈과 두 가지 일관성 손실(전반적 구조 및 국소 패치 수준)을 도입함으로써 의미 기반 지도의 고해상도 텍스처 생성을 가능하게 하여, 복잡한 다중 카테고리 구멍에서 경계 선명도와 텍스처 타당성 측면에서 최신 기술을 크게 능가한다.

ABSTRACT

Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we introduce coherence priors between the semantics and textures which make it possible to concentrate on completing separate textures in a semantic-wise manner. Specifically, we adopt a multi-scale joint optimization framework to first model the coherence priors and then accordingly interleavingly optimize image inpainting and semantic segmentation in a coarse-to-fine manner. A Semantic-Wise Attention Propagation (SWAP) module is devised to refine completed image textures across scales by exploring non-local semantic coherence, which effectively mitigates mix-up of textures. We also propose two coherence losses to constrain the consistency between the semantics and the inpainted image in terms of the overall structure and detailed textures. Experimental results demonstrate the superiority of our proposed method for challenging cases with complex holes.

연구 동기 및 목표

  • 다중 의미 카테고리와 혼합 텍스처를 포함한 복잡한 구멍에서의 이미지 보정 과제를 해결하기 위해.
  • 의미와 텍스처 간 상호 일관성을 모델링하여 텍스처의 현실성과 경계 선명도를 향상시키기 위해.
  • 의미 인식형 비국소 주의를 통해 보정 과정을 지도함으로써 텍스처 혼동을 감소시키기 위해.
  • 다중 해상도, 군집에서 세분화로의 정밀 조정을 통해 의미 분할과 이미지 보정을 공동 최적화하기 위해.
  • 예측된 의미와 보정된 이미지 간 전반적 구조적 및 국소 패치 수준의 일관성을 강제하는 손실 함수를 개발하기 위해.

제안 방법

  • 다중 해상도 공동 최적화 프레임워크를 도입하여 군집에서 세분화로의 방식으로 이미지 보정과 의미 분할을 번갈아가며 정밀 조정한다.
  • 동일한 예측된 의미 클래스를 가진 영역에서만 텍스처 특징를 선택하는 의미 기반 주의 전파(SWAP) 모듈을 제안한다. 이는 관련 없는 텍스처 메우기를 감소시킨다.
  • 보정된 이미지와 예측된 분할 맵 간의 구조적 일치를 강제하기 위해 이미지 수준의 구조 일관성 손실을 설계한다.
  • 유사한 의미 카테고리의 알려진 패치 분포와 일치하도록 생성된 텍스처를 장려하기 위해 비국소 패치 수준의 일관성 손실을 도입한다.
  • 다양한 해상도에서 의미와 텍스처 간 상호작용을 모델링하기 위해 공유된 특징 표현을 사용한다.
  • 보정과 분할의 공동 최적화를 위해 작업별 헤드를 갖춘 공유 인코더-디코더 백본을 활용한다.
Figure 1: Upper part: Mapping between image textures and edges/semantics (Dot arrow - extraction of edges/semantics; solid arrow - texture generation). Notice that two similar edge patches in a green circle could be mapped to completely different semantic textures, but one semantic will be clearly m
Figure 1: Upper part: Mapping between image textures and edges/semantics (Dot arrow - extraction of edges/semantics; solid arrow - texture generation). Notice that two similar edge patches in a green circle could be mapped to completely different semantic textures, but one semantic will be clearly m

실험 결과

연구 질문

  • RQ1의미와 텍스처 간의 일관성 사전 지식이 다중 의미 영역을 포함한 복잡한 구멍에서의 이미지 보정 품질 향상에 기여하는가?
  • RQ2전반적 또는 의미 없는 주의보다 의미 기반 주의 전파가 텍스처 혼동을 줄이는가?
  • RQ3보정과 분할의 공동 최적화가 순차적 또는 독립적 학습보다 더 높은 성능을 내는가?
  • RQ4제안된 전반적 및 국소 일관성 손실이 구조적 및 텍스처 일관성을 유지하는 데 얼마나 효과적인가?
  • RQ5의미 주석이 없는 Places2와 같은 새로운 데이터셋에 대해 모델이 일반화 가능한가?

주요 결과

  • Outdoor Scenes에서 제안된 방법은 PSNR 21.18, SSIM 0.81, FID 38.15를 기록하여 베이스라인 및 아블레이션 변형을 능가한다.
  • SWAP 모듈만으로도 기준 모델 대비 PSNR 3.75점 향상되고 FID 18.85점 감소하여 텍스처 혼동 감소에 효과적임을 입증한다.
  • SWAP와 일관성 손실을 모두 추가한 Ours Full은 Ours (SWAP) 대비 PSNR +1.6점 향상되고 FID -1.21점 감소하여 최고 성능 기록.
  • Outdoor Scenes에서 평균 교차율(mIoU)은 0.71, Cityscapes에서는 0.57을 기록하여 SPG-Net과 SGE-Net을 초월하는 분할 정확도 확보.
  • Places2 데이터셋에서 의미 주석이 없음에도 불구하고, 일관성 손실이 강력한 시나리오 사전 지식을 제공하여 타당한 텍스처와 구조를 생성한다.
  • 시각적 비교 결과 EdgeConnect는 경계 모호성으로 인해 잘못된 텍스처를 생성하는 반면, 제안된 방법은 의미 일관성에 의해 더 사진적 현실감 있는 결과를 생성한다.
Figure 2: Proposed network architecture. At each scale, both the inpainted image and segmentation map are output from two task-specific heads to control the predicted structures of the shared features. SWAP is added between scales to progressively optimize the texture details of contextual feature.
Figure 2: Proposed network architecture. At each scale, both the inpainted image and segmentation map are output from two task-specific heads to control the predicted structures of the shared features. SWAP is added between scales to progressively optimize the texture details of contextual feature.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.