[논문 리뷰] Asymmetric GANs for Image-to-Image Translation
이 논문은 이미지 간 번역 작업에서 번역 및 복원 작업에 각각 다른 크기와 아키텍처를 가진 비대칭 생성기 두 개를 사용하는 새로운 GAN 프레임워크인 AsymmetricGAN을 제안한다. 작업에 특화된 요구사항에 맞게 네트워크 설계를 분리함으로써 AsymmetricGAN은 더 낮은 모델 용량으로도 향상된 이미지 품질과 안정성을 달성하며, 슈퍼바이즈드 및 언슈퍼바이즈드 번역 작업 모두에서 StarGAN과 GestureGAN과 같은 대칭 모델들을 능가한다.
Existing models for unsupervised image translation with Generative Adversarial Networks (GANs) can learn the mapping from the source domain to the target domain using a cycle-consistency loss. However, these methods always adopt a symmetric network architecture to learn both forward and backward cycles. Because of the task complexity and cycle input difference between the source and target domains, the inequality in bidirectional forward-backward cycle translations is significant and the amount of information between two domains is different. In this paper, we analyze the limitation of existing symmetric GANs in asymmetric translation tasks, and propose an AsymmetricGAN model with both translation and reconstruction generators of unequal sizes and different parameter-sharing strategy to adapt to the asymmetric need in both unsupervised and supervised image translation tasks. Moreover, the training stage of existing methods has the common problem of model collapse that degrades the quality of the generated images, thus we explore different optimization losses for better training of AsymmetricGAN, making image translation with higher consistency and better stability. Extensive experiments on both supervised and unsupervised generative tasks with 8 datasets show that AsymmetricGAN achieves superior model capacity and better generation performance compared with existing GANs. To the best of our knowledge, we are the first to investigate the asymmetric GAN structure on both unsupervised and supervised image translation tasks.
연구 동기 및 목표
- 번역과 복원의 복잡도와 정보 흐름이 다를 수 있는 비대칭 이미지 간 번역 작업에서 대칭 GAN의 한계를 해결하기 위해.
- 번역 및 복원 작업에 특화된 전용 생성기를 설계하여 모델 용량을 줄이고 생성 품질을 향상시키기 위해.
- 비대칭 구성 요소에 맞는 최적화 손실을 적용하여 훈련의 불안정성과 모드 붕괴를 완화하기 위해.
- 감독 및 비감독 이미지 간 번역 벤치마크 전반에서 효과성을 입증하기 위해.
제안 방법
- 다른 크기의 생성기 두 개를 도입: 더 큰 번역 생성기 $G^t$ 는 소스 도메인에서 타겟 도메인으로 매핑하기 위한 것이고, 더 작은 복원 생성기 $G^r$ 는 번역된 출력에서 원본 이미지를 복원하기 위한 것이다.
- $G^t$ 와 $G^r$ 의 첫 번째 합성곱층에서 필터 수를 다르게 설정(64 대비 4)하여 비대칭 파라미터 공유 및 아키텍처 특화를 가능하게 한다.
- 도메인 전용 가이던스(예: 레이블 $z_x$, $z_y$ 또는 손 뼈대 $l_x$, $l_y$)를 사용하여 번역 및 복원 과정을 조건화한다.
- 훈련 안정성 향상과 품질 향상을 위해 적대적, 사이클 일致성, 인지적 손실 등의 다중 손실 함수를 함께 최적화한다.
- 두 개의 스트림 아키텍처를 채택하여 $G^t$ 와 $G^r$ 은 별도로 훈련되지만 사이클 일치성을 유지하기 위해 함께 최적화된다.
- 모델의 비대칭성을 고려한 파라미터 공유 전략을 적용하여 대칭 GAN의 일률적 설계를 피한다.
실험 결과
연구 질문
- RQ1감독 및 비감독 설정 모두에서 비대칭 생성기 아키텍처가 대칭 대비 이미지 간 번역 성능을 향상시키는가?
- RQ2번역 및 복원 작업 간 모델 용량을 분리하면 더 높은 생성 품질과 더 작은 모델 크기를 달성하는가?
- RQ3제안된 비대칭 설계가 GAN에서 모드 붕괴를 어떻게 완화하고 훈련 안정성을 향상시키는가?
- RQ4AsymmetricGAN이 손 동작 간 번역 작업에서 StarGAN 및 GestureGAN과 같은 최신 기술을 얼마나 능가하는가?
- RQ5비대칭 프레임워크는 손 동작 번역을 넘어서 다른 이미지 번역 작업으로 일반화될 수 있는가?
주요 결과
- NTU Hand Digit 데이터셋에서 AsymmetricGAN은 복원 브랜치의 파라미터 수가 크게 줄었음에도 불구하고 PSNR 32.6686을 기록하여 SymmetricGAN(32.5740)과 GestureGAN(32.6091)을 모두 능가한다.
- Senz3D 데이터셋에서 AsymmetricGAN은 PSNR 31.5624와 AMT 점수 28.1을 기록하여 PG2, SAMG, DPIG, PoseGAN, GestureGAN 등 모든 베이스라인을 초월한다.
- NTU Hand Digit에서 FID 점수 6.7132, Senz3D에서 12.4326을 기록하여 베이스라인 대비 더 뛰어난 이미지 품질과 분포 일치를 보였다.
- AsymmetricGAN은 총 모델 파라미터 수를 11.434M(비대칭 GAN의 22.776M 대비)로 줄였음에도 불구하고 PSNR와 AMT를 향상시켜 효율성과 효과성을 동시에 입증했다.
- 제거 실험 결과 비대칭 설계가 복원 생성기의 파라미터 수를 극적으로 줄여도 성능 향상을 이끌어내는 것으로 확인되었다.
- 정성적 결과에서는 AsymmetricGAN이 특히 복잡한 동작 번역 작업에서 모든 비교 방법보다 더 사실적인 이미지와 더 섬세한 디테일을 생성하는 것으로 나타났다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.