Skip to main content
QUICK REVIEW

[論文レビュー] FiG-NeRF: Figure-Ground Neural Radiance Fields for 3D Object Category Modelling

Christopher Xie, Keunhong Park|arXiv (Cornell University)|Apr 17, 2021
3D Shape Modeling and Analysis参考文献 44被引用数 6
ひとこと要約

FiG-NeRFは、制約のない背景を有する日常的・非制御的な画像から、3次元オブジェクトカテゴリの再構築と前面・背景の分離を同時に学習する2成分のニューラルレイトランスフィールド(NeRF)モデルを提案する。前面を可変形状のNeRF、背景を幾何学的に固定されながら外観が変化するNeRFとしてモデル化することで、シルエットや2次元アノテーションを一切必要とせずに、最先端のビュー合成と非模様セグメンテーションを達成する。

ABSTRACT

We investigate the use of Neural Radiance Fields (NeRF) to learn high quality 3D object category models from collections of input images. In contrast to previous work, we are able to do this whilst simultaneously separating foreground objects from their varying backgrounds. We achieve this via a 2-component NeRF model, FiG-NeRF, that prefers explanation of the scene as a geometrically constant background and a deformable foreground that represents the object category. We show that this method can learn accurate 3D object category models using only photometric supervision and casually captured images of the objects. Additionally, our 2-part decomposition allows the model to perform accurate and crisp amodal segmentation. We quantitatively evaluate our method with view synthesis and image fidelity metrics, using synthetic, lab-captured, and in-the-wild data. Our results demonstrate convincing 3D object category modelling that exceed the performance of existing methods.

研究の動機と目的

  • 制約のない背景を有する日常的・非制御的な画像から、高品質な3次元オブジェクトカテゴリモデルを学習すること。
  • 2次元セグメンテーションマスクやシルエットに依存せずに、3次元再構築と前面・背景分離を同時に達成すること。
  • 2成分NeRFモデルの内在的分解を活用して、非模様セグメンテーションを実現すること。
  • 光度的監督のみを用いて、オブジェクトロンデータセットのような野生のデータに対して一般化できることを示すこと。
  • 合成データ、ラボで撮影されたデータ、および現実世界のデータにおいて、ビュー合成とインスタンス補間の両面で既存手法を上回ること。

提案手法

  • 二重NeRFアーキテクチャを採用:前面オブジェクトに可変形状NeRF、背景に幾何学的に固定されたNeRFをそれぞれ使用する。
  • 背景NeRFに幾何学的一定性を課して分離を促進し、背景構造(例:テーブル、顔)が視点間で安定していると仮定する。
  • スパarsity事前分布と、トップ-kハード例マイニングを用いた学習可能なベータ損失を適用し、明確なオブジェクト境界を促進する。
  • 前面および背景モデルの空間特徴に10周波数の位置符号化を適用し、カップの場合は幾何的複雑性が低いことから4周波数に制限する。
  • 前面および背景成分の両方について、体積レンダリング積分を数値的台形則で近似する。
  • 初期学習段階でランダムな密度摂動を導入し、モード崩壊を回避し、両成分のバランスの取れた学習を促進する。
Figure 1 : Overview of our system. We take as input a collection of RGB captures of scenes with objects of a category. Our method jointly learns to decompose the scenes into foreground and background (without supervision) and a 3D object category model, that enables applications such as instance int
Figure 1 : Overview of our system. We take as input a collection of RGB captures of scenes with objects of a category. Our method jointly learns to decompose the scenes into foreground and background (without supervision) and a 3D object category model, that enables applications such as instance int

実験結果

リサーチクエスチョン

  • RQ12成分NeRFモデルは、制約のない野生の画像から、3次元オブジェクトカテゴリの再構築と前面・背景分離を同時に学習できるか?
  • RQ2背景の幾何学的固定を維持しつつ外観変化を許容することで、エンドツーエンドNeRFと比較して3次元再構築およびセグメンテーション品質が向上するか?
  • RQ32次元監督や真値シルエットを一切使用せずに、高精細なビュー合成と非模様セグメンテーションを達成できるか?
  • RQ43次元監督を用いるベースラインと比較して、多様で動的な背景を持つ現実世界データにおいて、本手法の性能はいかがなっているか?
  • RQ5図形と背景への分解が、インスタンス補間やビュー外挿などのタスクにおいて、どの程度優れた一般化性能を提供するか?

主な発見

  • FiG-NeRFは、合成データおよび現実世界データの両方で、最先端の新規ビュー合成性能を達成しており、オブジェクトロンベンチマークでも同様に優れた結果を示している。
  • 追加の学習画像を一切使用しなくても、マスクR-CNNや専用のマットイングネットワークよりも明確で正確な非模様セグメンテーションを生成する。
  • 2成分アーキテクチャにより、多様な背景において一貫した前面・背景分離が可能であり、2次元アノテーションの必要がない。
  • ビュー合成の結果、すべてのデータセットで非可変形状NeRFおよびSRNベースラインと比較してPSNRおよびLPIPS指標で顕著な向上を示している。
  • 野生のデータに対して良好な一般化性能を示し、変動する背景を持つ日常的・非制御的な動画からも、3次元オブジェクトカテゴリを正常に再構築できる。
  • アブレーションスタディにより、背景固定仮定とスパarsity事前分布が、安定的かつ正確な分解に不可欠であることが確認された。
Figure 2 : Example setups for Glasses (top) and Cups (bottom) datasets. For the lab-captured Glasses dataset [ 20 ] , the background (left) is a mannequin, and each scene (right) is a different pair of glasses placed on the mannequin. For Cups , we build this from the Objectron [ 2 ] dataset of crow
Figure 2 : Example setups for Glasses (top) and Cups (bottom) datasets. For the lab-captured Glasses dataset [ 20 ] , the background (left) is a mannequin, and each scene (right) is a different pair of glasses placed on the mannequin. For Cups , we build this from the Objectron [ 2 ] dataset of crow

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。