Skip to main content
QUICK REVIEW

[논문 리뷰] Unbalanced Feature Transport for Exemplar-based Image Translation

Fangneng Zhan, Yingchen Yu|arXiv (Cornell University)|2021. 06. 19.
Generative Adversarial Networks and Image Synthesis인용 수 7
한 줄 요약

이 논문은 조건부 입력과 스타일 예시 간의 특징을 정렬하기 위해 비균형 최적 운반(UNBALANCED OPTIMAL TRANSPORT, UOT)을 사용하는 새로운 예시 기반 이미지 번역 프레임워크인 UNITE를 제안한다. 이는 높은 정밀도와 현실적인 이미지 생성을 가능하게 하며, 스타일 전달을 충실하게 유지한다. 적응형 질량 학습과 다단계 특징 운반을 도입함으로써 UNITE는 다대일 매칭 문제를 완화하고 세밀한 질감을 유지하며, 정량적·정성적 지표에서 최신 기법들을 능가한다.

ABSTRACT

Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge maps, generating high-fidelity realistic images with reference styles remains a grand challenge in conditional image-to-image translation. This paper presents a general image translation framework that incorporates optimal transport for feature alignment between conditional inputs and style exemplars in image translation. The introduction of optimal transport mitigates the constraint of many-to-one feature matching significantly while building up accurate semantic correspondences between conditional inputs and exemplars. We design a novel unbalanced optimal transport to address the transport between features with deviational distributions which exists widely between conditional inputs and exemplars. In addition, we design a semantic-activation normalization scheme that injects style features of exemplars into the image translation process successfully. Extensive experiments over multiple image translation tasks show that our method achieves superior image translation qualitatively and quantitatively as compared with the state-of-the-art.

연구 동기 및 목표

  • 참조 예시를 사용한 조건부 이미지 간 번역에서 충실한 스타일 제어 문제를 해결하기 위해.
  • 일반적으로 다대일 매칭과 세부 사항 손실을 유도하는 코사인 유사도 기반 특징 매칭의 한계를 극복하기 위해.
  • 조건부 입력과 예시 간 분포가 상이한 특징을 효과적으로 정렬하기 위한 방법을 설계하기 위해.
  • 다중 척도, 다단계 특징 운반을 통해 복잡한 질감을 유지하기 위해.
  • 이미지 생성 중에 의미 일관성을 유지하는 스타일 주입 메커니즘을 개발하기 위해.

제안 방법

  • 조건부 입력과 예시 간 특징 정렬을 위해 비균형 최적 운반(UOT)을 도입하고, 분포 불일치를 다루기 위해 적응형 질량 학습을 적용한다.
  • 번역 과정에서 다양한 척도의 세밀한 질감을 유지하기 위해 다단계 특징 운반 전략을 활용한다.
  • 정렬된 스타일 특징을 번역 네트워크에 주입하면서 의미 일관성을 유지하기 위해 의미 활성화 정규화(SEACE) 레이어를 설계한다.
  • 특징 운반 및 번역 네트워크를 종합적으로 엔드 투 엔드로 최적화하여 정밀도와 스타일 정확도를 향상시킨다.
  • 코사인 유사도에서 사용하는 개별 특징 매칭의 함정을 피하기 위해 특징을 전체적으로 매칭하기 위해 최적 운반을 활용한다.
  • 이중 스트림 아키텍처를 활용한다: 특징 정렬을 위한 특징 운반 네트워크와 이미지 합성을 위한 번역 네트워크.
Figure 1: Different feature matching in image translation: Cosine Similarity tends to match each feature separately which often leads to many-to-one matching. Classical optimal transport ( Classical OT ) suppresses the many-to-one matching problem but it matches all feature points including undesire
Figure 1: Different feature matching in image translation: Cosine Similarity tends to match each feature separately which often leads to many-to-one matching. Classical optimal transport ( Classical OT ) suppresses the many-to-one matching problem but it matches all feature points including undesire

실험 결과

연구 질문

  • RQ1조건부 입력과 예시 간 특징 정렬을 어떻게 향상시켜 다대일 매칭을 줄이고 세밀한 세부 정보를 유지할 수 있는가?
  • RQ2비균형 최적 운반은 조건부 입력과 예시 간 분포 불일치를 다루는 데 어떤 역할을 하는가?
  • RQ3다단계 특징 운반은 이미지 번역에서 복잡한 질감을 유지하는 데 어떻게 기여하는가?
  • RQ4새로운 정규화 기법은 의미 일관성을 유지하면서 효과적으로 스타일 특징을 주입할 수 있는가?
  • RQ5제안된 방법은 현실성과 스타일 충실도 측면에서 기존의 예시 기반 이미지 번역 접근법보다 어느 정도 뛰어나게 성능을 발휘하는가?

주요 결과

  • UNITE는 CelebA-HQ에서 프리셰트 인ception 거리(FID)를 13.15로 기록하여 기준 모델인 SPADE(31.50) 및 기타 최신 기법들을 크게 앞서며 뛰어난 성능을 보였다.
  • 스윈 거리(SWD)를 14.91로 감소시켜 생성된 이미지와 진짜 이미지 간 분포 유사도가 향상되었음을 시사한다.
  • LPIPS 점수는 0.213을 기록하여 실제 이미지와 높은 지각 유사도를 확보하였으며, 사용자 설문 조사에서 30%의 선호도를 기록하였다.
  • 제거 실험을 통해 UOT에 적응형 질량 학습을 통합함으로써 잘못된 매칭과 다대일 매칭이 감소하여 기존 최적 운반 기법 대비 성능 향상이 확인되었다.
  • 다단계 운반 및 SEACE 구성 요소는 정성적 제거 실험을 통해 질감 유지와 스타일 일관성 향상에 기여하는 것으로 입증되었다.
  • UNITE는 CelebA-HQ 및 DeepFashion을 포함한 여러 데이터셋에서 뛰어난 다양성과 현실성을 보이며, 예시에 대한 충실한 스타일 전달을 실현하였다.
Figure 2: The framework of our proposed network: The Conditional Input and Exemplar are fed to feature extractors $F_{X}$ and $F_{Z}$ to extract feature vectors $X$ and $Z$ . The mass (or weight) of the feature vectors ( $\alpha$ and $\beta$ masses) are then determined collectively by $X$ and $Z$ .
Figure 2: The framework of our proposed network: The Conditional Input and Exemplar are fed to feature extractors $F_{X}$ and $F_{Z}$ to extract feature vectors $X$ and $Z$ . The mass (or weight) of the feature vectors ( $\alpha$ and $\beta$ masses) are then determined collectively by $X$ and $Z$ .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.