Skip to main content
QUICK REVIEW

[논문 리뷰] Positional Encoding as Spatial Inductive Bias in GANs

Rui Xu, Xintao Wang|arXiv (Cornell University)|2020. 12. 09.
Generative Adversarial Networks and Image Synthesis참고 문헌 46인용 수 7
한 줄 요약

이 논문은 컨volution GAN 생성자에서의 제로 패딩이 의도치 않게 비균형적인 암묵적 위치 인코딩을 생성함으로써, 이동 불변성에도 불구하고 고해상도 이미지 생성을 가능하게 한다는 점을 드러낸다. 공간적 인덕티브 바이어스를 향상시키기 위해 저자들은 명시적 위치 인코딩—카르테시안 공간 그리드와 2D 사인파 인코딩—을 제안하며, 이를 통해 새로운 다중 해상도 학습 전략(MS-PIE)을 가능하게 하여, 단일 256×256 StyleGAN2가 1024×1024까지 고품질 이미지를 생성할 수 있도록 하고, SinGAN의 이미지 조작 능력도 크게 향상시킨다.

ABSTRACT

SinGAN shows impressive capability in learning internal patch distribution despite its limited effective receptive field. We are interested in knowing how such a translation-invariant convolutional generator could capture the global structure with just a spatially i.i.d. input. In this work, taking SinGAN and StyleGAN2 as examples, we show that such capability, to a large extent, is brought by the implicit positional encoding when using zero padding in the generators. Such positional encoding is indispensable for generating images with high fidelity. The same phenomenon is observed in other generative architectures such as DCGAN and PGGAN. We further show that zero padding leads to an unbalanced spatial bias with a vague relation between locations. To offer a better spatial inductive bias, we investigate alternative positional encodings and analyze their effects. Based on a more flexible positional encoding explicitly, we propose a new multi-scale training strategy and demonstrate its effectiveness in the state-of-the-art unconditional generator StyleGAN2. Besides, the explicit spatial inductive bias substantially improve SinGAN for more versatile image manipulation.

연구 동기 및 목표

  • SinGAN과 StyleGAN2와 같은 GAN에서 이동 불변성 컨볼루션 생성자가 공간적으로 i.i.d. 입력을 받음에도 불구하고 구조적이고 고해상도 이미지를 생성할 수 있는 이유를 탐구하는 것.
  • 제로 패딩이 유도하는 의도치 않은 공간적 편향과 그가 특징 맵 전반에 걸쳐 비균형적으로 분포하는 방식을 분석하는 것.
  • 암묵적 제로패딩 편향보다 우월한 대안으로, 카르테시안 공간 그리드와 2D 사인파 인코딩을 포함한 명시적 위치 인코딩을 제안하는 것.
  • 단일 생성자가 여러 해상도에서 이미지를 합성할 수 있도록 하는 다중 해상도 학습 전략(MS-PIE)을 개발하는 것.
  • 명시적 공간적 인덕티브 바이어스를 활용해 SinGAN의 고해상도 이미지 외삽 및 조작 작업에서의 강건성과 유연성을 향상시키는 것.

제안 방법

  • 저자들은 제로 패딩이 비대칭 경계 효과로 인해 암묵적 위치 인코딩을 유도함을 밝히며, 이는 특징 맵의 가장자리에서 중심으로 점차 퍼져나가는 경향이 있음을 확인한다.
  • 두 가지 명시적 위치 인코딩—정규화된 카르테시안 공간 그리드와 2D 사인파 인코딩—을 평가하며, 둘 다 균형 잡힌 공간적 인덕티브 바이어스를 제공하도록 설계되었다.
  • 다양한 입력 해상도에서 학습 중 위치 인코딩을 적응시키는 Multi-Scale training with PositIon Encodings(MS-PIE)라는 새로운 다중 해상도 학습 전략을 제안한다.
  • 이 방법은 생성자의 입력에 명시적 위치 인코딩을 통합하여 제로 패딩에서 유래한 암묵적 편향을 대체하거나 보완한다.
  • StyleGAN2와 SinGAN에서 이 방법을 검증하였으며, 패딩 없는 설정과 다양한 스케일 변환 조건에서의 분석도 수행하였다.
  • 채널 수를 절반으로 줄인 경량 버전의 생성자도 테스트하여, 명시적 위치 인코딩을 사용할 경우 성능 저하 없이도 성능 유지가 가능함을 입증하였다.
Figure 1: Images sampled from the internal patch distribution learned by SinGAN. Above the dotted line, we present sampled balloons with standard SinGAN and padding-free SinGAN. A more challenging case of generating a school of fish is shown below the dotted line. (c)-(f) show the effects of differe
Figure 1: Images sampled from the internal patch distribution learned by SinGAN. Above the dotted line, we present sampled balloons with standard SinGAN and padding-free SinGAN. A more challenging case of generating a school of fish is shown below the dotted line. (c)-(f) show the effects of differe

실험 결과

연구 질문

  • RQ1이동 불변성 컨볼루션 생성자가 공간적으로 i.i.d. 노이즈 입력을 받음에도 불구하고 구조적이고 고해상도 이미지를 생성할 수 있는 이유는 무엇인가요?
  • RQ2제로 패딩은 어떻게 암묵적이고 비균형적인 위치 인코딩을 유도하며, GAN 생성자 내 특징 맵 분포에 영향을 미치는가요?
  • RQ3카르테시안 공간 그리드와 2D 사인파 인코딩과 같은 명시적 위치 인코딩이 암묵적 제로패딩 편향보다 더 균형 잡히고 효과적인 공간적 인덕티브 바이어스를 제공할 수 있는가요?
  • RQ4명시적 위치 인코딩을 활용해 사전 학습된 단일 256×256 StyleGAN2 생성자를 다중 해상도 이미지 합성에 적응시킬 수 있는가요?
  • RQ5명시적 위치 인코딩은 고해상도 이미지 외삽 및 조작 작업에서 SinGAN의 강건성과 유연성을 어떻게 향상시키는가요?

주요 결과

  • 제로 패딩은 GAN 생성자에 의도치 않게 암묵적 위치 인코딩을 유도하며, 이는 이동 불변성에도 불구하고 구조적이고 고해상도 이미지 생성을 가능하게 한다.
  • 제로 패딩에서 유래한 암묵적 편향은 비균형적이며, 이미지 가장자리와 모서리에서 더 강한 영향을 미치며 중심 영역의 성능 저하를 초래한다. 이는 SinGAN의 물고기 생성 예시에서 관찰되었다.
  • 명시적 위치 인코딩—카르테시안 공간 그리드와 2D 사인파 인코딩—은 이미지 전반에 걸쳐 더 일관되고 안정적인 공간적 구조를 생성하여 패치 재현성과 전반적인 일관성을 향상시킨다.
  • 제안된 MS-PIE 전략을 통해 단일 256×256 StyleGAN2 생성자는 512×512, 896×256, 1024×1024 해상도에서 고품질 이미지를 생성할 수 있었으며, FID 및 P&R 점수에서 기준 방법 대비 뚜렷한 향상이 관찰되었다.
  • 명시적 위치 인코딩을 적용한 SinGAN은 더 나은 구조 보존과 스케일 업 시 왜곡 감소를 통해 뛰어난 이미지 조작 성능을 달성하였으며, 제로패딩 기반 기준 대비 뛰어난 성능을 보였다.
  • MS-PIE 생성자의 경량 버전은 채널 너비를 절반으로 줄여도 경쟁 가능한 성능을 유지하며, 명시적 위치 인코딩의 효율성과 강건성을 입증하였다.
Figure 2 : Illustration for the convolutional procedure with zero padding. We move the padding in each layer to the input feature and regard the whole convolutional network as a convolutional layer with a large kernel.
Figure 2 : Illustration for the convolutional procedure with zero padding. We move the padding in each layer to the input feature and regard the whole convolutional network as a convolutional layer with a large kernel.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.