Skip to main content
QUICK REVIEW

[論文レビュー] GeoDream: Disentangling 2D and Geometric Priors for High-Fidelity and Consistent 3D Generation

Baorui Ma, Haoge Deng|arXiv (Cornell University)|Nov 29, 2023
Computer Graphics and Visualization Techniques被引用数 6
ひとこと要約

GeoDreamは、明示的な3D幾何的事前知識と事前学習済み2D拡散モデルの事前知識を組み合わせる分離型フレームワークを導入し、テキストプロンプトから高精細で3D整合性を持つテクスチャ付きメッシュを生成する。複数視点の2D予測からコストボリュームを構築することで強固な3D事前知識を獲得し、分離型最適化により繰り返し精緻化することで、ジェイナス問題やアーチファクトを顕著に低減する一方で、1024×1024解像度の高精細なリアルリズムと意味的整合性を維持する。

ABSTRACT

Text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models has shown great promise but still suffers from inconsistent 3D geometric structures (Janus problems) and severe artifacts. The aforementioned problems mainly stem from 2D diffusion models lacking 3D awareness during the lifting. In this work, we present GeoDream, a novel method that incorporates explicit generalized 3D priors with 2D diffusion priors to enhance the capability of obtaining unambiguous 3D consistent geometric structures without sacrificing diversity or fidelity. Specifically, we first utilize a multi-view diffusion model to generate posed images and then construct cost volume from the predicted image, which serves as native 3D geometric priors, ensuring spatial consistency in 3D space. Subsequently, we further propose to harness 3D geometric priors to unlock the great potential of 3D awareness in 2D diffusion priors via a disentangled design. Notably, disentangling 2D and 3D priors allows us to refine 3D geometric priors further. We justify that the refined 3D geometric priors aid in the 3D-aware capability of 2D diffusion priors, which in turn provides superior guidance for the refinement of 3D geometric priors. Our numerical and visual comparisons demonstrate that GeoDream generates more 3D consistent textured meshes with high-resolution realistic renderings (i.e., 1024 $ imes$ 1024) and adheres more closely to semantic coherence.

研究の動機と目的

  • 2D拡散モデルを用いたテキストから3Dへの生成において、持続的な不整合な3D幾何的構造(ジェイナス問題)と深刻なアーチファクトを解消すること。
  • 高精細性、多様性、現実的でない3D出力を犠牲にすることなく、3D整合性を向上させること。
  • 侵襲的なファインチューニングを回避しつつ、明示的な3D事前知識を通じて2D拡散モデルに3D認識能力を付与すること。
  • コストボリュームアグリゲーションを用いて、不完全な複数視点2D予測から強固な3D事前知識を構築すること。
  • 2Dと3D事前知識の間で双方向の精緻化ループを確立し、幾何的・意味的整合性を向上させること。

提案手法

  • 分散と投影操作を用いて、複数視点の2D画像特徴をコストボリュームに集約することで3D幾何的事前知識を構築する。
  • 3D特徴ネットワークを用いてコストボリュームを処理し、個々の視点の不一致に依存しない空間的に整合性のある3D事前知識を生成する。
  • LoRAを適用した2D拡散モデルを訓練することで2Dと3D事前知識を分離し、ベースモデルのファインチューニングを伴わずに3D事前知識を活用する。
  • 2D拡散事前知識が3D事前知識の精緻化を繰り返し最適化によって誘導するフィードバックループを用いて、3D幾何的事前知識を精緻化する。
  • DMTetベースのメッシュ抽出パイプラインを用いて、最適化された3D表現から高解像度のテクスチャ付きメッシュを生成する。
  • 複数のカメラビューにわたる生成された3D出力を、事前学習済み2D拡散モデルと整合させるために重み付きノイズ除去損失を用いる。
Figure 1 : GeoDream alleviates the Janus problems by incorporating explicit 3D priors with 2D diffusion priors. GeoDream generates consistent multi-view rendered images and rich details textured meshes. We remove rendering background to achieve a clearer visualization.
Figure 1 : GeoDream alleviates the Janus problems by incorporating explicit 3D priors with 2D diffusion priors. GeoDream generates consistent multi-view rendered images and rich details textured meshes. We remove rendering background to achieve a clearer visualization.

実験結果

リサーチクエスチョン

  • RQ1複数視点の2D予測から導出される明示的な3D幾何的事前知識は、テキストから3Dへの生成における3D整合性を向上させることができるか?
  • RQ22Dと3D事前知識を分離することで、ファインチューニングなしに2D拡散モデルがより良い3D認識行動を示せるか?
  • RQ3コストボリュームアグリゲーションは、複数視点2D予測からの不一致を効果的にフィルタリングし、強固な3D事前知識を生成できるか?
  • RQ42Dと3D事前知識間の双方向精緻化は、3D生成における幾何的・意味的整合性をどのように向上させるか?
  • RQ5GeoDreamは、Zero123 や MVDream などの多様な複数視点拡散モデルにも一般化できるか?

主な発見

  • GeoDreamは1024×1024解像度で高精細なテクスチャを備えた3Dアセットを生成し、ジェイナス問題やアーチファクトを顕著に低減する。
  • 本手法は、非対称的で想像力豊かなオブジェクトを含む多様で複雑なプロンプトにおいても、優れた3D整合性と意味的整合性を達成する。
  • アブレーションスタディにより、Zero123 や Zero123++ などの異なる複数視点拡散モデルから生成されたソースビューに対しても、GeoDreamが良好に一般化することが確認された。
  • 2Dと3D事前知識の間の分離型精緻化ループは、エンドツーエンドまたは共同最適化ベースラインと比較して、より優れた幾何的構造とテクスチャの詳細を実現する。
  • コストボリュームに基づく3D事前知識構築法は、複数視点予測からの不一致を効果的に緩和し、直接的な複数視点整合性に依存する手法を上回る性能を発揮する。
  • 定量的3Dメトリクスと定性的比較により、GeoDreamは先行手法と比較してより現実的で幾何的に整合性のあるメッシュを生成することが示された。
Figure 2 : The overview of GeoDream. (a) 3D priors training. (b) Incorporating 3D priors with 2D diffusion priors.
Figure 2 : The overview of GeoDream. (a) 3D priors training. (b) Incorporating 3D priors with 2D diffusion priors.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。