[论文解读] Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era
在 AIGC 时代对文本到3D方法的综合综述,涵盖3D数据表示、基础技术、CLIP和扩散先验,以及3D应用与挑战。
Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions and AI-generated content (AIGC). Thanks to advancements in text-to-image and 3D modeling technologies, like neural radiance field (NeRF), text-to-3D has emerged as a nascent yet highly active research field. Our work conducts a comprehensive survey on this topic and follows up on subsequent research progress in the overall field, aiming to help readers interested in this direction quickly catch up with its rapid development. First, we introduce 3D data representations, including both Structured and non-Structured data. Building on this pre-requisite, we introduce various core technologies to achieve satisfactory text-to-3D results. Additionally, we present mainstream baselines and research directions in recent text-to-3D technology, including fidelity, efficiency, consistency, controllability, diversity, and applicability. Furthermore, we summarize the usage of text-to-3D technology in various applications, including avatar generation, texture generation, scene generation and 3D editing. Finally, we discuss the agenda for the future development of text-to-3D.
研究动机与目标
- 解释3D数据表示(欧几里德与非欧几里德)及其在文本到3D中的相关性。
- 评述支撑文本到3D的基础技术,包括 NeRF、CLIP 和扩散模型。
- 总结知名的文本到3D模型及其如何利用图文先验进行3D生成。
- 综述文本到3D的应用(头像、纹理、场景、形状变换)及实际挑战。
- 讨论文本到3D中的保真度、速度、一致性、可控性和适用性问题。
提出的方法
- 将3D数据表示分为欧几里德(体素网格、多视图图像)和非欧几里德(网格、点云、神经场)两类。
- 描述基于 NeRF 的场景表示及用于3D合成的可微分渲染。
- 解释基于 CLIP 的文本-图像先验及其在引导3D生成中的作用。
- 总结基于扩散模型的先验以及它们如何与 NeRF 或显式3D表示集成。
- 评述开创性与后续的文本到3D方法(DreamFusion、Magic3D、CLIP-NeRF、CLIP-Mesh、DreamFields、3D-CLFusion、CLIP-Sculptor 等)。
- 概述从文本提示到通过先验与优化或潜在扩散得到的优化3D输出的典型流程。

实验结果
研究问题
- RQ1在文本到3D任务中,哪些3D数据表示最有效或最实用?
- RQ2NeRF、CLIP 和扩散先验如何促进文本引导的3D生成?
- RQ3文本到3D的主要架构与数据挑战有哪些,近来工作如何应对数据稀缺?
- RQ4文本到3D在头像、纹理、场景和形状变换中的主要应用与局限性是什么?
- RQ5当前文本到3D方法中出现的保真度、速度、一致性和可控性问题以及可能的缓解方法?
主要发现
- 文本到3D依赖于欧几里德与非欧几里德表示的混合,神经场(例如 NeRF、SDF)使输出更加灵活且高分辨率。
- 预训练的文本到图像扩散模型和 CLIP 提供强先验,使在有限的3D数据下实现文本引导的3D合成成为可能。
- 开创性方法(DreamFusion、Magic3D、CLIP-NeRF、DreamFields)演示了利用图文先验从文本提示优化3D内容的可行性,通常通过基于 NeRF 的表示。
- 后续工作聚焦于加速生成、提升3D一致性,以及实现多类和可编辑的3D内容(如 3D-CLFusion、CLIP-Sculptor、3DFuse、CompoNeRF)。
- 应用涵盖文本引导的3D头像、纹理和场景生成,以及形状变换,面临持续的保真度、速度和语义一致性挑战。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。