[논문 리뷰] A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
이 설문은 AIGC를 정의하고 GANs에서 ChatGPT까지의 역사를 추적하며 단일 모달 및 다중 모달 생성 모델, 기초 이론, 응용 및 열린 도전 과제를 분석한다.
Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its impressive performance. In fact, ChatGPT and other Generative AI (GAI) techniques belong to the category of Artificial Intelligence Generated Content (AIGC), which involves the creation of digital content, such as images, music, and natural language, through AI models. The goal of AIGC is to make the content creation process more efficient and accessible, allowing for the production of high-quality content at a faster pace. AIGC is achieved by extracting and understanding intent information from instructions provided by human, and generating the content according to its knowledge and the intent information. In recent years, large-scale models have become increasingly important in AIGC as they provide better intent extraction and thus, improved generation results. With the growth of data and the size of the models, the distribution that the model can learn becomes more comprehensive and closer to reality, leading to more realistic and high-quality content generation. This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimodal interaction and multimodal interaction. From the perspective of unimodality, we introduce the generation tasks and relative models of text and image. From the perspective of multimodality, we introduce the cross-application between the modalities mentioned above. Finally, we discuss the existing open problems and future challenges in AIGC.
연구 동기 및 목표
- AI-generated content (AIGC)와 AI 강화 생성 프로세스에 대한 공식적인 정의와 철저한 조사를 제공한다.
- 비전 및 언어 모달리티 전반에 걸친 생성 모델의 역사적 발전을 검토한다.
- 현대 GAI 시스템에서 사용되는 기초 기술과 구성요소를 요약한다.
- 단일 모달 및 다중 모달 생성의 발전과 그들의 아키텍처적 경향을 분석한다.
- AIGC의 응용, 도전과제, 위험 및 향후 방향에 대해 논의한다.
제안 방법
- CV, NLP, VL 맥락에서 생성 모델의 역사를 검토한다.
- 기초 기술(트랜스포머, 사전학습된 언어 모델, 확산, GANs, VAE, 흐름 모델)을 요약한다.
- RLHF 및 인간 선호도에 대한 출력 일치를 위한 역할을 설명한다.
- 생성 모델을 단일 모달과 다중 모달로 분류하고 핵심 아키텍처(GPT-유사 디코더, 인코더-디코더, 비전-언어 모델)를 종합한다.
- 대규모 모델을 가능하게 하는 하드웨어, 분산 학습, 클라우드 컴퓨팅에 대해 논의한다.
- AIGC의 미해결 문제 및 향후 연구 방향을 파악한다.
실험 결과
연구 질문
- RQ1AIGC가 무엇이며 생성 파이프라인 내에서 Formal 정의는 무엇인가?
- RQ2AIGC를 가능하게 하는 기초 기술은 무엇이며 어떻게 발전해 왔는가(트랜스포머, 확산, GANs, VAE 등)?
- RQ3단일 모달 대 다중 모달 생성 모델의 주요 진보와 아키텍처 경향은 무엇인가?
- RQ4AIGC와 관련된 주요 응용 분야 및 열린 문제(위험, 평가, 정렬, 거버넌스)는 무엇인가?
- RQ5향후 방향은 AIGC의 개발과 사회적 영향에 어떤 방향으로 형성될 가능성이 있는가?
주요 결과
- AIGC는 인간이 제공한 의도에 의해 안내되는 AI 모델이 생성한 컨텐츠로 정의되며, 더 큰 데이터, 모델, 컴퓨트에 의해 개선된다.
- 트랜스포머와 대규모 사전학습 언어 모델이 텍스트 및 비전 작업 전반의 현대적 생성을 뒷받침한다.
- RLHF 및 관련 정렬 방법은 AIGC 출력의 유용성 및 안전성을 향상시키는 데 중심적이다.
- 다중 모달 모델(예: 비전-언어, 텍스트-코드)은 교차 모달 생성 및 프롬프트 기능을 가능하게 한다.
- 단일 모달에서 다중 모달 통합으로의 경로가 더 풍부한 콘텐츠 생성과 더 넓은 응용을 가능하게 한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.