Skip to main content
QUICK REVIEW

[論文レビュー] A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Yihan Cao, Siyu Li|arXiv (Cornell University)|Mar 7, 2023
Topic Modeling被引用数 398
ひとこと要約

本調査はAIGCを定義し、GANsからChatGPTまでの歴史をたどり、単一モーダルおよび多モーダル生成モデル・基礎・応用・未解決の課題を分析する。

ABSTRACT

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its impressive performance. In fact, ChatGPT and other Generative AI (GAI) techniques belong to the category of Artificial Intelligence Generated Content (AIGC), which involves the creation of digital content, such as images, music, and natural language, through AI models. The goal of AIGC is to make the content creation process more efficient and accessible, allowing for the production of high-quality content at a faster pace. AIGC is achieved by extracting and understanding intent information from instructions provided by human, and generating the content according to its knowledge and the intent information. In recent years, large-scale models have become increasingly important in AIGC as they provide better intent extraction and thus, improved generation results. With the growth of data and the size of the models, the distribution that the model can learn becomes more comprehensive and closer to reality, leading to more realistic and high-quality content generation. This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimodal interaction and multimodal interaction. From the perspective of unimodality, we introduce the generation tasks and relative models of text and image. From the perspective of multimodality, we introduce the cross-application between the modalities mentioned above. Finally, we discuss the existing open problems and future challenges in AIGC.

研究の動機と目的

  • AI生成コンテンツ(AIGC)とAI強化生成プロセスの正式な定義と徹底的な調査を提供する。
  • 視覚と言語のモダリティ全体にわたる生成モデルの歴史的発展をレビューする。
  • 現代のGAIシステムで用いられる基盤技術と構成要素を概説する。
  • 単一モーダル生成と多モーダル生成の進展とそのアーキテクチャ動向を分析する。
  • AIGCの応用、課題、リスク、および将来の方向性について論じる。

提案手法

  • CV、NLP、およびVLの文脈における生成モデルの歴史をレビューする。
  • 基盤技術(transformers、事前学習済み言語モデル、拡散モデル、GAN、VAE、フローモデル)を要約する。
  • RLHFと人間の嗜好に出力を合わせる際の役割を説明する。
  • 生成モデルを単一モーダルと多モーダルに分類し、主要なアーキテクチャ(GPT風デコーダ、エンコーダ-デコーダ、ビジョン-言語モデル)を統合する。
  • 大規模モデルを可能にする計算ハードウェア、分散学習、およびクラウドコンピューティングを議論する。
  • AIGCにおける未解決の問題と将来の研究方向を特定する。

実験結果

リサーチクエスチョン

  • RQ1AI Generated Content(AIGC)とは何か、生成パイプライン内で正式にどのように定義されるか?
  • RQ2AIGCを可能にする基盤技術は何か、それらはどのように進化してきたか(transformers、拡散、GAN、VAEなど)?
  • RQ3単一モーダル生成モデルと多モーダル生成モデルの主要な進展と、それらのアーキテクチャ動向は何か?
  • RQ4AIGCに関連する主要な応用と未解決の課題(リスク、評価、整合、ガバナンス)は何か?
  • RQ5AIGCの発展と社会的影響を形作ると考えられる将来の方向性は何か?

主な発見

  • AIモデルによって人間が提供する意図に導かれて生成されるコンテンツをAIGCと定義し、データ・モデル・計算リソースの大規模化によって改善が進む。
  • Transformersと大規模事前学習済み言語モデルは、テキストおよびビジョンタスク全体の現代的生成を支える。
  • RLHFおよび関連する整合手法は、AIGC出力の有用性と安全性を高める上で中心的である。
  • マルチモーダルモデル(例:視覚-言語、テキスト-コード)は、モーダル間の生成と prompting 能力を可能にする。
  • 単一モーダルから多モーダル統合への推移が、より豊かなコンテンツ生成とより広範な応用を可能にしている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。