Skip to main content
QUICK REVIEW

[논문 리뷰] FiG-NeRF: Figure-Ground Neural Radiance Fields for 3D Object Category Modelling

Christopher Xie, Keunhong Park|arXiv (Cornell University)|2021. 04. 17.
3D Shape Modeling and Analysis참고 문헌 44인용 수 6
한 줄 요약

FiG-NeRF는 자유로운 배경을 가진 일상적인 이미지에서 3D 객체 카테고리 재구성과 전경-배경 분리를 동시에 학습하는 이중 구성 요소 신경 렌더링 필드(NeRF) 모델을 제안한다. 전경을 변형 가능한 NeRF로, 배경을 기하학적으로 고정된, 외관이 변하는 NeRF로 모델링함으로써, 2D 세그멘테이션 마스크나 실루엣 정보 없이도 최신 기술 수준의 시각 합성 및 모달 세그멘테이션을 달성한다.

ABSTRACT

We investigate the use of Neural Radiance Fields (NeRF) to learn high quality 3D object category models from collections of input images. In contrast to previous work, we are able to do this whilst simultaneously separating foreground objects from their varying backgrounds. We achieve this via a 2-component NeRF model, FiG-NeRF, that prefers explanation of the scene as a geometrically constant background and a deformable foreground that represents the object category. We show that this method can learn accurate 3D object category models using only photometric supervision and casually captured images of the objects. Additionally, our 2-part decomposition allows the model to perform accurate and crisp amodal segmentation. We quantitatively evaluate our method with view synthesis and image fidelity metrics, using synthetic, lab-captured, and in-the-wild data. Our results demonstrate convincing 3D object category modelling that exceed the performance of existing methods.

연구 동기 및 목표

  • 자유로운 배경을 가진 일상적인 이미지에서 고품질의 3D 객체 카테고리 모델을 학습하는 것.
  • 2D 세그멘테이션 마스크나 실루엣 정보에 의존하지 않고도 3D 재구성과 전경-배경 분리를 동시에 달성하는 것.
  • 이중 구성 요소 NeRF 모델의 내재된 분해 특성을 활용해 모달 세그멘테이션을 가능하게 하는 것.
  • 광학적 감독만을 사용하여 Objectron 데이터셋과 같은 실세계 데이터에 대해 일반화 성능을 입증하는 것.
  • 합성, 실험실 촬영, 실세계 데이터에서의 시각 합성 및 인스턴스 보간 성능에서 기존 방법들을 능가하는 것.

제안 방법

  • 이중-NeRF 아키텍처를 사용: 전경 객체에 대해 변형 가능한 NeRF와 배경에 대해 기하학적으로 고정된 NeRF를 각각 적용.
  • 배경 NeRF에 기하학적 일관성 제약 조건을 도입하여 분리를 유도함. 배경의 구조(예: 테이블, 얼굴 등)가 시야 간에 안정적이라고 가정.
  • 희소성 사전과 학습 가능한 베타 손실을 사용하며, 상위-k 하드 예외 마이닝을 통해 날카운 전경 경계를 촉진.
  • 전경 및 배경 모델의 공간적 특징에 대해 10개의 주파수를 사용한 위치 인코딩을 적용하며, 기하학적 복잡도가 낮은 컵의 경우 4개의 주파수를 사용.
  • 전경 및 배경 구성 요소 모두에 대해 볼륨 렌더링 적분을 수치적 적분으로 근사.
  • 초기 학습 단계에서 무작위 밀도 변형을 도입하여 모드 붕괴를 방지하고 양 구성 요소의 균형 잡힌 학습을 촉진.
Figure 1 : Overview of our system. We take as input a collection of RGB captures of scenes with objects of a category. Our method jointly learns to decompose the scenes into foreground and background (without supervision) and a 3D object category model, that enables applications such as instance int
Figure 1 : Overview of our system. We take as input a collection of RGB captures of scenes with objects of a category. Our method jointly learns to decompose the scenes into foreground and background (without supervision) and a 3D object category model, that enables applications such as instance int

실험 결과

연구 질문

  • RQ1이중 구성 요소 NeRF 모델은 자유로운, 실세계의 이미지에서 3D 객체 카테고리 재구성과 전경-배경 분리를 동시에 학습할 수 있는가?
  • RQ2기하학적 구조를 고정하고 외관 변화만 허용함으로써, 종단 간 NeRF와 비교해 3D 재구성 및 세그멘테이션 품질을 향상시킬 수 있는가?
  • RQ32D 감독이나 참값 실루엣 없이도 고해상도의 시각 합성과 모달 세그멘테이션을 달성할 수 있는가?
  • RQ43D 감독을 사용하는 기존 기반 모델들과 비교해, 다양한 동적 배경을 가진 실세계 데이터에서 이 방법의 성능은 어떠한가?
  • RQ5도형과 배경으로의 분해가 인스턴스 보간 및 시각 외삽과 같은 작업에서 더 나은 일반화 성능을 제공하는 데에 얼마나 기여하는가?

주요 결과

  • FiG-NeRF는 합성 및 실세계 데이터셋, 특히 Objectron 벤치마크에서 최신 기술 수준의 새로운 시각 합성 성능를 달성한다.
  • 추가 학습 이미지 없이도 Mask R-CNN와 전용 매트팅 네트워크보다 더 날카우며 정확한 모달 세그멘테이션을 생성한다.
  • 이중 구성 요소 아키텍처 덕분에 다양한 배경에서 일관된 전경-배경 분리를 달성하며, 2D 애너테이션의 필요 없이도 가능하다.
  • 시각 합성 결과에서 모든 데이터셋에서 비변형 NeRF 및 SRN 기반 모델 대비 PSNR 및 LPIPS 지표에서 뚜렷한 향상을 보였다.
  • 모델은 실세계 데이터에 대해 잘 일반화되며, 다양한 배경을 가진 일상적인 영상에서 3D 객체 카테고리를 성공적으로 재구성한다.
  • 제거 실험을 통해 기하학적 고정 가정과 희소성 사전가 모델의 안정적이고 정확한 분해에 필수적임을 확인했다.
Figure 2 : Example setups for Glasses (top) and Cups (bottom) datasets. For the lab-captured Glasses dataset [ 20 ] , the background (left) is a mannequin, and each scene (right) is a different pair of glasses placed on the mannequin. For Cups , we build this from the Objectron [ 2 ] dataset of crow
Figure 2 : Example setups for Glasses (top) and Cups (bottom) datasets. For the lab-captured Glasses dataset [ 20 ] , the background (left) is a mannequin, and each scene (right) is a different pair of glasses placed on the mannequin. For Cups , we build this from the Objectron [ 2 ] dataset of crow

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.