Skip to main content
QUICK REVIEW

[論文レビュー] MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction

Zehao Yu, Songyou Peng|arXiv (Cornell University)|Jun 1, 2022
Advanced Vision and Imaging被引用数 166
ひとこと要約

MonoSDF は monocular depth と normal cues をニューラル・インプリシット表面再構成に統合し、多様なシーンと表現(MLP および グリッドベース)で精度と収束を改善します。

ABSTRACT

In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural networks. State-of-the-art neural implicit methods allow for high-quality reconstructions of simple scenes from many input views. Yet, their performance drops significantly for larger and more complex scenes and scenes captured from sparse viewpoints. This is caused primarily by the inherent ambiguity in the RGB reconstruction loss that does not provide enough constraints, in particular in less-observed and textureless areas. Motivated by recent advances in the area of monocular geometry prediction, we systematically explore the utility these cues provide for improving neural implicit surface reconstruction. We demonstrate that depth and normal cues, predicted by general-purpose monocular estimators, significantly improve reconstruction quality and optimization time. Further, we analyse and investigate multiple design choices for representing neural implicit surfaces, ranging from monolithic MLP models over single-grid to multi-resolution grid representations. We observe that geometric monocular priors improve performance both for small-scale single-object as well as large-scale multi-object scenes, independent of the choice of representation.

研究の動機と目的

  • マルチビュー情報が限られているまたはテクスチャが欠如している場合の堅牢な3D再構成を動機づける。
  • モノキュラ深度と法線予測がニューラル・インプリシット表面をどのように制約できるかを調査する。
  • ニューラル・インプリシット表現(MLP、密なグリッド、単一/多解像度グリッド)を体系的に比較する。
  • オブジェクトレベルと室/シーンスケールのデータセットで MonoSDF を評価し、スケーラビリティと効率を評価する。

提案手法

  • シーンの形状を、密な SDF グリッド、単一の MLP、単一解像度の特徴グリッド、そして多解像度特徴グリッドの4つのオプションでパラメータ化された SDF として表現する。
  • 最適化のために微分可能なボリュームレンダリングを用いて RGB、深度、および法線を取得する。
  • 事前学習済みの Omnidata モデルを用いてモノキュラの手掛かり(深度と法線)を予測し、それらを監視信号として使用する。
  • RGB再構成、Eikonal正則化、深度整合性、法線整合性を、バッチごとのスケール/シフト整列とともに組み合わせた損失を定義する。
  • キューを強化した損失を用いて、Adam でシーン表現と外観を共同最適化する。

実験結果

リサーチクエスチョン

  • RQ1テクスチャが少ないまたは複数物体のシーンで、モノキュラ深度と法線 cues はニューラル・インプリシット表面再構成を改善できるか。
  • RQ2どのニューラル・インプリシット表現(MLP 対 グリッドベース)が品質と収束速度の点でモノキュラ priors から最も恩恵を受けるか。
  • RQ3深度と法線 priors はRGBベースの監督と異なるシーンスケールでどのように相互作用し、補完するか。

主な発見

  • モノキュラ手掛かりは、MLPと多解像度グリッドの両方の再構成品質を顕著に向上させ、深度と法線の両方を用いると最良の結果になる。
  • Replica において手掛かりなしのグリッドベース表現では多解像度グリッドが他より優れるが、手掛かりがあると両方のアーキテクチャが改善され、収束速度が向上する。
  • モノキュラ手掛かりを用いたMLPは、多くの厳しい設定で全体的に最も強い結果を達成する一方、グリッドより収束は遅い。
  • ScanNet では MLP バリアントがベースラインを上回り、より滑らかで詳細な再構成を生成する。
  • DTU のスパースビュー実験は、手掛かりが MLP とグリッドの双方の性能を高めることを示し、密なビューではグリッドが手掛かりの恩恵をより受ける。
  • multimodal 評価は、モノキュラの事前情報がオブジェクトレベル・室レベル・シーンレベルのデータ全体でスケーラブルで堅牢な再構成を可能にすることを示す。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。