Skip to main content
QUICK REVIEW

[Paper Review] GAN Computers Generate Arts? A Survey on Visual Arts, Music, and Literary Text Generation using Generative Adversarial Network

Sakib Shahriar|arXiv (Cornell University)|Aug 9, 2021
Generative Adversarial Networks and Image Synthesis4 citations
TL;DR

This survey explores the application of Generative Adversarial Networks (GANs) in generating visual arts, music, and literary text, providing a comprehensive review of GAN architectures, performance comparisons, and challenges in creative AI. It highlights key advancements in photorealistic image synthesis, conditional generation, and text-to-image modeling, while identifying limitations and future research directions in artistic content generation using deep generative models.

ABSTRACT

"Art is the lie that enables us to realize the truth." - Pablo Picasso. For centuries, humans have dedicated themselves to producing arts to convey their imagination. The advancement in technology and deep learning in particular, has caught the attention of many researchers trying to investigate whether art generation is possible by computers and algorithms. Using generative adversarial networks (GANs), applications such as synthesizing photorealistic human faces and creating captions automatically from images were realized. This survey takes a comprehensive look at the recent works using GANs for generating visual arts, music, and literary text. A performance comparison and description of the various GAN architecture are also presented. Finally, some of the key challenges in art generation using GANs are highlighted along with recommendations for future work.

Motivation & Objective

  • To comprehensively review recent advancements in using GANs for generating visual arts, music, and literary text.
  • To analyze and compare various GAN architectures across different creative domains.
  • To identify key challenges in generating high-quality, coherent, and meaningful artistic content using GANs.
  • To provide recommendations for future research in artistic generative modeling using deep learning.

Proposed method

  • Systematic review of peer-reviewed literature on GAN-based art generation from 2015 to 2021.
  • Categorization of GAN models based on architecture, loss functions, and training strategies in visual, audio, and textual generation.
  • Performance comparison of GAN variants such as StyleGAN, BigGAN, and conditional GANs across metrics like FID, IS, and human evaluation.
  • Analysis of attention mechanisms, self-attention modules, and progressive growing techniques in improving generation quality.
  • Incorporation of conditional inputs (e.g., class labels, text embeddings) to guide content generation in a disentangled manner.
  • Evaluation of both quantitative metrics and qualitative human assessments in assessing artistic output fidelity and creativity.

Experimental results

Research questions

  • RQ1How have GANs been adapted to generate high-fidelity visual arts, including photorealistic images and artistic styles?
  • RQ2What architectural innovations in GANs have enabled improved generation of structured and diverse musical compositions?
  • RQ3How effective are GAN-based models in generating coherent and contextually relevant literary text, including story generation and captioning?
  • RQ4What are the major limitations in current GAN-based art generation, particularly in terms of mode collapse, training instability, and semantic consistency?
  • RQ5What future research directions can enhance the creativity, controllability, and interpretability of GAN-generated artistic content?

Key findings

  • StyleGAN and its variants achieved state-of-the-art results in generating high-resolution, diverse, and photorealistic human faces with disentangled attributes.
  • Conditional GANs and their variants enabled significant progress in text-to-image generation, allowing precise control over generated content using textual prompts.
  • BigGAN demonstrated the scalability of GANs in generating high-fidelity images across multiple classes, achieving low Fréchet Inception Distance (FID) scores.
  • In music generation, GANs with latent space modeling and waveform-based architectures showed promise in generating coherent and rhythmically structured compositions.
  • Despite progress, GANs still face challenges such as mode collapse, poor generalization to rare concepts, and difficulty in maintaining long-range semantic coherence in text generation.
  • Human evaluation remains essential, as quantitative metrics like FID and Inception Score often fail to correlate with perceived artistic quality or creativity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.