[Paper Review] A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
This survey defines AIGC, traces its history from GANs to ChatGPT, and analyzes unimodal and multimodal generative models, foundations, applications, and open challenges.
Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its impressive performance. In fact, ChatGPT and other Generative AI (GAI) techniques belong to the category of Artificial Intelligence Generated Content (AIGC), which involves the creation of digital content, such as images, music, and natural language, through AI models. The goal of AIGC is to make the content creation process more efficient and accessible, allowing for the production of high-quality content at a faster pace. AIGC is achieved by extracting and understanding intent information from instructions provided by human, and generating the content according to its knowledge and the intent information. In recent years, large-scale models have become increasingly important in AIGC as they provide better intent extraction and thus, improved generation results. With the growth of data and the size of the models, the distribution that the model can learn becomes more comprehensive and closer to reality, leading to more realistic and high-quality content generation. This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimodal interaction and multimodal interaction. From the perspective of unimodality, we introduce the generation tasks and relative models of text and image. From the perspective of multimodality, we introduce the cross-application between the modalities mentioned above. Finally, we discuss the existing open problems and future challenges in AIGC.
Motivation & Objective
- Provide a formal definition and thorough survey of AI-generated content (AIGC) and the AI-enhanced generation process.
- Review the historical development of generative models across vision and language modalities.
- Summarize foundational techniques and components used in modern GAI systems.
- Analyze advances in unimodal and multimodal generation and their architectural trends.
- Discuss applications, challenges, risks, and future directions for AIGC.
Proposed method
- Review the history of generative models in CV, NLP, and VL contexts.
- Summarize foundation techniques (transformers, pre-trained language models, diffusion, GANs, VAEs, flow models).
- Explain RLHF and its role in aligning outputs with human preferences.
- Categorize generative models into unimodal and multimodal and synthesize key architectures (GPT-like decoders, encoder-decoders, vision-language models).
- Discuss computing hardware, distributed training, and cloud computing enabling large-scale models.
- Identify open problems and future research directions in AIGC.
Experimental results
Research questions
- RQ1What is AI-Generated Content (AIGC) and how is it formally defined within the generation pipeline?
- RQ2What are the foundational technologies enabling AIGC and how have they evolved (transformers, diffusion, GANs, VAEs, etc.)?
- RQ3What are the major advances in unimodal versus multimodal generative models and their architectural trends?
- RQ4What are the primary applications and open challenges (risks, evaluation, alignment, governance) associated with AIGC?
- RQ5What future directions are likely to shape the development of AIGC and its societal impact?
Key findings
- AIGC is defined as content generated by AI models guided by human-provided intent, with improvements driven by larger data, models, and compute.
- Transformers and large pre-trained language models underpin modern generation across text and vision tasks.
- RLHF and related alignment methods are central to improving usefulness and safety of AIGC outputs.
- Multimodal models (e.g., vision-language, text-code) enable cross-modal generation and prompting capabilities.
- There is a trajectory from unimodal to multimodal integration enabling richer content generation and broader applications.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.