[Paper Review] A Comprehensive Survey on 3D Content Generation
This survey presents a novel taxonomy for 3D content generation, categorizing methods into 3D native, 2D prior-based, and hybrid approaches. It reviews 60 key papers, identifies critical challenges in quality, controllability, speed, and data scarcity, and outlines future directions including foundation models and improved benchmarks for AIGC-3D.
Recent years have witnessed remarkable advances in artificial intelligence generated content(AIGC), with diverse input modalities, e.g., text, image, video, audio and 3D. The 3D is the most close visual modality to real-world 3D environment and carries enormous knowledge. The 3D content generation shows both academic and practical values while also presenting formidable technical challenges. This review aims to consolidate developments within the burgeoning domain of 3D content generation. Specifically, a new taxonomy is proposed that categorizes existing approaches into three types: 3D native generative methods, 2D prior-based 3D generative methods, and hybrid 3D generative methods. The survey covers approximately 60 papers spanning the major techniques. Besides, we discuss limitations of current 3D content generation techniques, and point out open challenges as well as promising directions for future work. Accompanied with this survey, we have established a project website where the resources on 3D content generation research are provided. The project page is available at https://github.com/hitcslj/Awesome-AIGC-3D.
Motivation & Objective
- To systematically categorize recent advances in 3D content generation through a new taxonomy of 3D native, 2D prior-based, and hybrid methods.
- To provide a comprehensive review of approximately 60 key papers across major 3D generation techniques.
- To identify unresolved challenges in geometry quality, texture detail, controllability, speed, and data scarcity.
- To propose future research directions, including foundation models for 3D generation and improved automated evaluation benchmarks.
- To support the research community by curating a public resource website with up-to-date materials on AIGC-3D.
Proposed method
- Proposes a new three-tier taxonomy: 3D native generative methods, 2D prior-based 3D generative methods, and hybrid 3D generative methods.
- Reviews 60 representative papers across the three categories, focusing on architectures, training paradigms, and input modalities.
- Analyzes 3D representation techniques, including explicit (point clouds, meshes, voxels) and implicit (NeRF, Gaussian Splatting, signed distance functions).
- Examines diffusion-based methods, including score distillation sampling (SDS) for NeRF optimization and multi-view diffusion priors.
- Introduces and discusses emerging pipelines for 4D (dynamic) content generation, such as static-to-dynamic and video-to-3D reconstruction.
- Evaluates current benchmarks and proposes improvements, including automated human-aligned evaluators and holistic metrics for geometric and textural fidelity.
Experimental results
Research questions
- RQ1How can 3D content generation methods be systematically categorized to reflect the current technological landscape?
- RQ2What are the core technical differences and trade-offs between 3D native, 2D prior-based, and hybrid 3D generation approaches?
- RQ3What are the primary limitations in current 3D generation in terms of geometry quality, texture detail, controllability, and inference speed?
- RQ4How can large-scale, diverse 3D datasets be collected and leveraged to improve unsupervised and self-supervised 3D generation?
- RQ5What role can foundation models and multimodal LLMs play in enabling programmatic 3D design and high-level control in future AIGC-3D systems?
Key findings
- The survey identifies that 2D prior-based methods, especially those leveraging pre-trained image diffusion models, have become dominant due to the scarcity of 3D data.
- 3D native generative methods face challenges due to limited availability of large-scale 3D datasets, despite their potential for direct 3D generation.
- Hybrid methods such as one2345++ and 4D-fy demonstrate improved quality by combining 2D priors with 3D diffusion or Gaussian Splatting optimization.
- Current 4D generation pipelines, including 4DGen and DreamGaussian4d, achieve improved temporal coherence and geometric fidelity using multi-view diffusion priors.
- Despite progress, generating high-fidelity, compact meshes with rich textures and accurate material properties remains a key unsolved challenge.
- Automated evaluation remains underdeveloped, with human evaluation still the gold standard, though tools like the Human-Aligned Evaluator (Wu et al., 2024) are emerging.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.