Skip to main content
QUICK REVIEW

[論文レビュー] Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era

Chenghao Li, Chaoning Zhang|arXiv (Cornell University)|May 10, 2023
Human Motion and Animation被引用数 24
ひとこと要約

AIGC時代のテキスト→3D手法の包括的な調査で、3Dデータ表現、基盤技術、CLIPと拡散の priors、そして3Dアプリケーションと課題を網羅している。

ABSTRACT

Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions and AI-generated content (AIGC). Thanks to advancements in text-to-image and 3D modeling technologies, like neural radiance field (NeRF), text-to-3D has emerged as a nascent yet highly active research field. Our work conducts a comprehensive survey on this topic and follows up on subsequent research progress in the overall field, aiming to help readers interested in this direction quickly catch up with its rapid development. First, we introduce 3D data representations, including both Structured and non-Structured data. Building on this pre-requisite, we introduce various core technologies to achieve satisfactory text-to-3D results. Additionally, we present mainstream baselines and research directions in recent text-to-3D technology, including fidelity, efficiency, consistency, controllability, diversity, and applicability. Furthermore, we summarize the usage of text-to-3D technology in various applications, including avatar generation, texture generation, scene generation and 3D editing. Finally, we discuss the agenda for the future development of text-to-3D.

研究の動機と目的

  • ユークリッド空間と非ユークリッド空間を含む3Dデータ表現を説明し、それらがテキストから3Dへのタスクにおいてどのように関連するかを明らかにする。
  • NeRF、CLIP、拡散モデルを含む、テキストから3Dを実現する基盤技術をレビューする。
  • 画像とテキストの事前知識を活用して3D生成を行う、顕著なテキスト→3Dモデルを要約する。
  • テキスト→3Dアプリケーション(アバター、テクスチャ、シーン、形状変換)と実践的な課題を調査する。
  • 忠実度、速度、一貫性、操作性、適用性の課題を議論する。

提案手法

  • 3Dデータ表現をユークリッド(ボクセルグリッド、マルチビュー画像)と非ユークリッド(メッシュ、点群、ニューラルフィールド)に分類する。
  • NeRFベースのシーン表現と3D合成のための微分可能レンダリングを説明する。
  • CLIPを用いたテキスト-画像の事前知識と、それが3D生成を指示する際の役割を説明する。
  • 拡散モデルに基づく事前知識と、それらがどのようにNeRFや明示的な3D表現と統合されるかを要約する。
  • 先駆的手法(DreamFusion、Magic3D、CLIP-NeRF、CLIP-Mesh、DreamFields、3D-CLFusion、CLIP-Sculptor など)とその後の派生手法をレビューする。
  • 事前知識と最適化または潜在拡散を介して、テキストプロンプトから最適化された3D出力へ至る典型的なパイプラインを概説する。
Figure 1. Voxel representation of Stanford bunny, the picture is obtained from (Shi et al . , 2022 ) .
Figure 1. Voxel representation of Stanford bunny, the picture is obtained from (Shi et al . , 2022 ) .

実験結果

リサーチクエスチョン

  • RQ1テキスト→3Dタスクに最も効果的または実用的な3Dデータ表現は何か?
  • RQ2NeRF、CLIP、拡散の事前知識は、テキスト指示による3D生成にどのように寄与するか?
  • RQ3テキスト→3Dの主なアーキテクチャ上およびデータ上の課題は何か、そして最近の研究はデータ不足にどう対処しているか?
  • RQ4アバター、テクスチャ、シーン、変換におけるテキスト→3Dの主な応用と制限は何か?
  • RQ5現在のテキスト→3D手法で生じる忠実度・速度・一貫性・制御性の課題は何で、どのように緩和できるか?

主な発見

  • テキスト→3Dはユークリッド表現と非ユークリッド表現の組み合わせに依存し、神経場(例:NeRF、SDF)が柔軟で高解像度の出力を可能にする。
  • 事前学習済みのテキスト→画像拡散モデルとCLIPは、限られた3Dデータでのテキスト誘導3D合成を可能にする強力な事前知識を提供する。
  • 先駆的手法は、画像とテキストの事前知識を用いてテキストプロンプトから3Dコンテンツを最適化する実現可能性を示しており、しばしNeRFベースの表現を介している。
  • フォローアップの研究は、生成の高速化、3Dの一貫性向上、マルチクラス対応および編集可能な3Dコンテンツの実現に焦点を当てる(例:3D-CLFusion、CLIP-Sculptor、3DFuse、CompoNeRF)。
  • アプリケーションは、テキスト誘導の3Dアバター、テクスチャ、シーン生成、形状変換を含み、忠実度・速度・意味的一貫性の課題が継続している。
Figure 2. Multi-view representation of Stanford bunny, the picture is obtained from (Park et al . , 2016 ) .
Figure 2. Multi-view representation of Stanford bunny, the picture is obtained from (Park et al . , 2016 ) .

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。