[论文解读] A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
本综述将 AIGC 定义为由 AI 模型在人工提供的意图指导下生成的内容,并追溯其从 GANs 到 ChatGPT 的历史,同时分析单模态与多模态生成模型、基础、应用及未解挑战。
Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its impressive performance. In fact, ChatGPT and other Generative AI (GAI) techniques belong to the category of Artificial Intelligence Generated Content (AIGC), which involves the creation of digital content, such as images, music, and natural language, through AI models. The goal of AIGC is to make the content creation process more efficient and accessible, allowing for the production of high-quality content at a faster pace. AIGC is achieved by extracting and understanding intent information from instructions provided by human, and generating the content according to its knowledge and the intent information. In recent years, large-scale models have become increasingly important in AIGC as they provide better intent extraction and thus, improved generation results. With the growth of data and the size of the models, the distribution that the model can learn becomes more comprehensive and closer to reality, leading to more realistic and high-quality content generation. This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimodal interaction and multimodal interaction. From the perspective of unimodality, we introduce the generation tasks and relative models of text and image. From the perspective of multimodality, we introduce the cross-application between the modalities mentioned above. Finally, we discuss the existing open problems and future challenges in AIGC.
研究动机与目标
- 提供对 AI 生成内容(AIGC)及 AI 增强生成过程的正式定义和全面综述。
- 回顾跨视觉与语言模态的生成模型的历史发展。
- 总结现代生成 AI 系统中使用的基础技术与组件。
- 分析单模态与多模态生成的发展及其架构趋势。
- 讨论 AIGC 的应用、挑战、风险及未来方向。
提出的方法
- 回顾在计算机视觉(CV)、自然语言处理(NLP)和视觉-语言(VL)背景下的生成模型历史。
- 总结基础技术(Transformers、预训练语言模型、扩散、GAN、VAE、流模型)。
- 解释 RLHF 及其在使输出与人类偏好对齐中的作用。
- 将生成模型分类为单模态和多模态,并综合关键架构(GPT 风格的解码器、编码器-解码器、视觉-语言模型)。
- 讨论推动大规模模型的计算硬件、分布式训练与云计算。
- 指出 AIGC 的开放性问题与未来研究方向。
实验结果
研究问题
- RQ1AI 生成内容(AIGC)是什么,及其在生成流程中的正式定义?
- RQ2推动 AIGC 的基础技术有哪些,它们如何演变(Transformers、扩散、GAN、VAE 等)?
- RQ3单模态与多模态生成模型的主要进展及其架构趋势有哪些?
- RQ4与 AIGC 相关的主要应用及开放挑战(风险、评估、对齐、治理)有哪些?
- RQ5哪些未来方向可能影响 AIGC 的发展及其对社会的影响?
主要发现
- AIGC 被定义为由 AI 模型在人工提供的意图指导下生成的内容,其改进由更大规模的数据、模型和计算能力推动。
- Transformers 和大型预训练语言模型支撑着跨文本与视觉任务的现代生成。
- RLHF 及相关对齐方法对于提高 AIGC 输出的有效性与安全性至关重要。
- 多模态模型(如视觉-语言、文本-代码)实现跨模态生成与提示能力。
- 存在从单模态到多模态整合的趋势,使内容生成更丰富、应用更广泛。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。