Skip to main content
QUICK REVIEW

[논문 리뷰] Semantic Scene Completion via Integrating Instances and Scene in-the-Loop

Yingjie Cai, Xuesong Chen|arXiv (Cornell University)|2021. 04. 08.
Advanced Vision and Imaging참고 문헌 43인용 수 7
한 줄 요약

이 논문은 3D 세분화 장면 완성도 향상을 위해 장면에서 인스턴스로, 인스턴스에서 장면으로의 반복적 정합을 반복하는 새로운 반복 프레임워크인 SISNet을 제안한다. 뚜렷한 장면 완성된 장면에서 인스턴스를 분리하고, 가중치 공유된 반복 루프를 통해 장면 및 인스턴스 수준의 사전 지식을 활용함으로써, SISNet은 NYU, NYUCAD, SUNCG-RGBD 데이터셋에서 최신 기술 수준(SOTA)의 성능을 달성하며, 세밀한 형태 복원을 향상시키고 카테고리 혼동을 줄였다.

ABSTRACT

Semantic Scene Completion aims at reconstructing a complete 3D scene with precise voxel-wise semantics from a single-view depth or RGBD image. It is a crucial but challenging problem for indoor scene understanding. In this work, we present a novel framework named Scene-Instance-Scene Network ( extit{SISNet}), which takes advantages of both instance and scene level semantic information. Our method is capable of inferring fine-grained shape details as well as nearby objects whose semantic categories are easily mixed-up. The key insight is that we decouple the instances from a coarsely completed semantic scene instead of a raw input image to guide the reconstruction of instances and the overall scene. SISNet conducts iterative scene-to-instance (SI) and instance-to-scene (IS) semantic completion. Specifically, the SI is able to encode objects' surrounding context for effectively decoupling instances from the scene and each instance could be voxelized into higher resolution to capture finer details. With IS, fine-grained instance information can be integrated back into the 3D scene and thus leads to more accurate semantic scene completion. Utilizing such an iterative mechanism, the scene and instance completion benefits each other to achieve higher completion accuracy. Extensively experiments show that our proposed method consistently outperforms state-of-the-art methods on both real NYU, NYUCAD and synthetic SUNCG-RGBD datasets. The code and the supplementary material will be available at \url{https://github.com/yjcaimeow/SISNet}.

연구 동기 및 목표

  • 단일 뷰 RGBD 입력에서의 불완전하고 모호한 3D 장면 이해 도전 과제를 해결하기 위해, 특히 세밀한 형태 세부 사항과 의미적으로 혼동이 쉬운 영역에 초점을 맞춘다.
  • 기존의 SSC 및 SIC 방법이 인스턴스 수준 제약을 무시하거나 장면 맥락을 활용하지 못해 인스턴스 탐지 및 완성도가 떨어지는 문제를 해결하기 위해 노력한다.
  • 장면 및 인스턴스 수준 간의 双방향 정보 흐름을 가능하게 하는 통합 프레임워크를 개발하여 전체 완성 정확도를 향상시키는 것을 목표로 한다.
  • 반복적이고 파rameter 효율적인 메커니즘을 통해 장면 및 인스턴스 수준의 의미 사전 지식을 통합함으로써 벤치마크 데이터셋에서 최신 기술 수준의 성능을 달성하는 것

제안 방법

  • 프레임워크는 다중 모odal 입력(2D 세분화 맵 및 TSDF)을 사용하여 초기 장면 $S_0$를 구성하는 것으로 시작한다.
  • 장면에서 인스턴스로의 정합(SI 정합)은 $S_0$ 내 객체를 국소화하고, 각 인스턴스를 더 높은 해상도로 볼록화하여 세밀한 3D 형태를 복원한다.
  • 인스턴스에서 장면으로의 정합(IS 정합)은 개선된 인스턴스 형태를 장면에 통합하여 맥락 피드백을 통해 전체 장면 완성도를 향상시킨다.
  • 장면 및 인스턴스 정합을 번갈아 가며 수행하는 반복적이고 가중치 공유 메커니즘을 활용하여 장면 및 인스턴스 예측 간 상호 강화를 가능하게 한다.
  • 다단계 감독을 통해 엔드 투 엔드로 훈련되는 반복 과정은 인스턴스 세부 사항과 전반적인 장면 의미를 점진적으로 개선할 수 있도록 한다.
  • 카테고리 사전 지식과 장면 맥락을 활용하여 객체 카테고리의 모호함(예: 창문 vs. 벽)을 해결하고 탐지 정확도를 향상시킨다.
Figure 2: Overview of the Proposed Method . SISNet consists of iterative scene-to-instance completion and instance-to-scene completion stages. Given single-view RGBD images, TSDF from the depth map and semantic volume from the reprojection of 2D semantic segmentation are input into the initial scene
Figure 2: Overview of the Proposed Method . SISNet consists of iterative scene-to-instance completion and instance-to-scene completion stages. Given single-view RGBD images, TSDF from the depth map and semantic volume from the reprojection of 2D semantic segmentation are input into the initial scene

실험 결과

연구 질문

  • RQ1粗안된 장면에 인스턴스 수준의 세부 정보를 통합하면 3D 세분화 장면 완성도 정확도가 향상되는가?
  • RQ2장면 및 인스턴스 수준 간의 반복적 정합은 단일 단계 또는 비반복적 접근보다 더 나은 성능을 내는가?
  • RQ3장면 맥락이 감독 신호로 사용될 경우, 인스턴스 탐지 및 형태 완성도에 상당한 영향을 미치는가?
  • RQ4복잡한, 가림이 있는, 또는 의미적으로 혼동이 쉬운 영역을 다룰 때 제안된 방법은 최신 기술 수준의 SSC 및 SIC 방법과 비교해 어떻게 성능을 내는가?
  • RQ5성능과 계산 비용의 균형을 고려할 때, 장면-인스턴스-장면 루프에서 최적의 반복 횟수는 얼마인가?

주요 결과

  • NYU 데이터셋에서 SISNet은 한 번의 반복으로 기준 $S_0$ 대비 SSC IoU에서 5.2% 향상되었고, 두 번의 반복으로 추가로 2% 향상되었다.
  • NYUCAD에서 메서드는 한 번의 반복으로 SSC IoU를 4.4% 향상시키고, 두 번의 반복으로 추가로 2% 향상시켜 데이터셋 간 일관된 성과 향상을 보였다.
  • SUNCG-RGBD에서 SISNet은 계산 비용이 유사한 것으로 비해 SC IoU에서 Sketch보다 8.1% 높고, SSC IoU에서 23.2% 높게 성과를 냈다.
  • 제거 실험 결과, 인스턴스 정합 단계를 생략하면 평균적으로 SSC 성능이 2.8% 저하됨을 확인하여, 이 단계가 개선 과정에서 핵심적인 역할을 한다는 것을 입증했다.
  • 초기 장면 완성 단계는 mAP 및 재현율에서 약 10% 향상시켜, 장면 맥락이 인스턴스 국소화에 도움이 된다는 것을 보여주었다.
  • 두 번 이상의 반복은 유의미한 성능 향상 없이, 두 번의 반복이 최적의 성능을 달성하는 데 충분함을 시사했다.
Figure 3: Architecture details of initial scene completion. Given TSDF and semantic volume, to infer missing invisible shape details, we exploit DDR to encode features and propagate information with skip layers, yielding a whole scene completion results.
Figure 3: Architecture details of initial scene completion. Given TSDF and semantic volume, to infer missing invisible shape details, we exploit DDR to encode features and propagate information with skip layers, yielding a whole scene completion results.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.