Skip to main content
QUICK REVIEW

[논문 리뷰] AesUST: Towards Aesthetic-Enhanced Universal Style Transfer

Zhizhong Wang, Zhanjie Zhang|arXiv (Cornell University)|2022. 08. 27.
Generative Adversarial Networks and Image Synthesis인용 수 4
한 줄 요약

AesUST는 전용 미적 감지기와 Aesthetic-aware Style-Attention (AesSA) 모듈을 통해 학습된 인간이 기쁨을 느끼는 시각적 미적 특징을 스타일 전이에 통합함으로써 시각적 조화와 현실감을 크게 향상시킨 혁신적인 미적 향상 통합 스타일 전이 프레임워크를 제안한다. 실험과 사용자 연구를 통해 AesUST는 실제 예술가가 그린 그림과 구분이 불가능한 결과를 생성하며, 인지적 품질과 미적 현실감 모두에서 기존 최고 성능(SOTA) 방법을 능가한다.

ABSTRACT

Recent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious patterns and evident artifacts, making the results easy to spot from real paintings. To address this limitation, we propose AesUST, a novel Aesthetic-enhanced Universal Style Transfer approach that can generate aesthetically more realistic and pleasing results for arbitrary styles. Specifically, our approach introduces an aesthetic discriminator to learn the universal human-delightful aesthetic features from a large corpus of artist-created paintings. Then, the aesthetic features are incorporated to enhance the style transfer process via a novel Aesthetic-aware Style-Attention (AesSA) module. Such an AesSA module enables our AesUST to efficiently and flexibly integrate the style patterns according to the global aesthetic channel distribution of the style image and the local semantic spatial distribution of the content image. Moreover, we also develop a new two-stage transfer training strategy with two aesthetic regularizations to train our model more effectively, further improving stylization performance. Extensive experiments and user studies demonstrate that our approach synthesizes aesthetically more harmonious and realistic results than state of the art, greatly narrowing the disparity with real artist-created paintings. Our code is available at https://github.com/EndyWon/AesUST.

연구 동기 및 목표

  • 스タイ링된 이미지에 조화가 맞지 않는 패atters와 잡음이 발생하는 통합 스타일 전이의 미적 현실감 결여 문제를 해결하기 위해.
  • 예술가가 그린 그림에서 유래한 통합된 인간이 기쁨을 느끼는 시각적 미적 특징을 통합하여 스타일링된 이미지의 현실감과 시각적 조화를 향상시키기 위해.
  • 콘텐츠 보존과 스타일 전이의 균형을 유지하면서도 미적 일관성을 확보하는 방법을 개발하기 위해.
  • 이중 미적 정규화를 적용한 이중 단계 전략을 사용해 모델을 효과적으로 훈련하기 위해.
  • AI가 생성한 스타일링된 이미지와 실제 예술가가 그린 예술작품 간의 인지적 격차를 좁히기 위해.

제안 방법

  • 대규모 예술가가 그린 그림 코퍼스를 기반으로 훈련된 미적 감지기를 도입하여 통합된 인간이 기쁨을 느끼는 시각적 미적 특징을 추출한다.
  • 스타일 이미지의 전반적인 미적 분포와 콘텐츠 이미지의 국소적 의미적 공간 분포를 기반으로 동적으로 스타일 패턴을 통합하는 Aesthetic-aware Style-Attention (AesSA) 모듈을 제안한다.
  • 이중 단계 훈련 전략을 적용한다: 제1단계는 콘텐츠 및 스타일 재구성으로 생성자 미리 훈련, 제2단계는 적대적 손실 및 미적 정규화로 보정 훈련.
  • 두 가지 미적 정규화를 적용한다: 하나는 콘텐츠 구조 보존과 스타일 전이 품질 유지를 위한 것이며, 다른 하나는 인지적 스타일 패턴의 정밀도 향상을 위한 것이다.
  • 실제 그림과 생성된 그림을 구별할 수 있도록 유도하기 위해 적대적 손실을 적용하여 현실적이며 잡음이 없는 출력을 장려한다.
  • 미적 감지기를 단순한 감지기 외에도 스타일 통합을 안내하는 특징 추출기로 활용한다.
Figure 2. Overview of our proposed AesUST. Note that we apply a two-stage transfer training strategy with a transfer learning fashion: (1) At stage I, the aesthetic discriminator only acts as the discriminator, and we pre-train the generator (AesSA module and decoder) using VGG features only. (2) At
Figure 2. Overview of our proposed AesUST. Note that we apply a two-stage transfer training strategy with a transfer learning fashion: (1) At stage I, the aesthetic discriminator only acts as the discriminator, and we pre-train the generator (AesSA module and decoder) using VGG features only. (2) At

실험 결과

연구 질문

  • RQ1학습된 미적 감지기가 예술가가 그린 그림에서 통합된 인간이 기쁨을 느끼는 특징을 효과적으로 캡처하여 스타일 전이를 안내할 수 있는가?
  • RQ2어떻게 하면 미적 특징을 스타일 전이 과정에 통합하여 시각적 조화와 현실감을 향상시킬 수 있는가?
  • RQ3제안된 이중 정규화를 적용한 이중 단계 훈련 전략이 종단간(end-to-end) 훈련보다 더 나은 스타일링 성능을 내는가?
  • RQ4AesUST는 얼마나 깊이 AI가 생성한 스타일링된 이미지와 실제 예술가가 그린 예술작품 간의 인지적 격차를 줄이는가?
  • RQ5AesSA 모듈은 스타일 전이 과정에서 전반적인 미적 일관성과 국소적 의미적 충실도 사이를 얼마나 민첩하게 균형을 맞출 수 있는가?

주요 결과

  • 사용자 연구에서 AesUST는 실제 그림(89.2%)에 가까운 87.4%의 위장성 점수를 기록하여 실제 예술작품과 거의 구별되지 않음을 시사한다.
  • A/B 선호도 테스트에서 AesUST는 68.7%의 표를 확보하여 AdaIN(52.3%)과 LST(55.1%)와 같은 SOTA 방법들을 크게 앞서며 사용자 선호도에서 압도적인 우위를 점한다.
  • 제거 실험 결과, 적대적 손실을 제거할 경우 심각한 잡음과 조화가 맞지 않는 패턴이 발생함을 확인하여, 이 손실이 현실감 확보에 핵심적인 역할을 함을 입증한다.
  • 첫 번째 미적 정규화(L_AR1)를 제거하면 콘텐츠 보존과 스타일 전이 품질이 모두 악화되며, 두 번째 정규화(L_AR2)를 제거하면 스타일 정밀도가 떨어짐을 확인하여 두 정규화가 필수적임을 증명한다.
  • 이중 단계 훈련 전략은 필수적이다: 제1단계의 미리 훈련은 효과적인 스타일 통합을 가능하게 하며, 제2단계의 미적 안내 보정 훈련은 최종 품질 향상에 기여한다.
  • AesUST는 512×512 이미지에서 약 25 FPS로 실시간으로 작동하여 AdaIN 및 LST와 같은 SOTA 방법과 동등한 효율성을 확보한다.
Figure 3. Aesthetic-aware style-attention (AesSA) module. $F_{c}$ , $F_{s}$ , and $F_{a}$ are content, style, and aesthetic features, respectively. “ $Norm$ ” denotes the mean-variance channel-wise normalization.
Figure 3. Aesthetic-aware style-attention (AesSA) module. $F_{c}$ , $F_{s}$ , and $F_{a}$ are content, style, and aesthetic features, respectively. “ $Norm$ ” denotes the mean-variance channel-wise normalization.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.