[Paper Review] Matterport3D: Learning from RGB-D Data in Indoor Environments
Matterport3D introduces a large-scale RGB-D dataset of 90 building-scale scenes with 194,400 RGB-D images and 10,800 panoramas, enabling diverse supervised and self-supervised indoor scene understanding tasks with precise global alignment and semantic annotations.
Access to large, diverse RGB-D datasets is critical for training RGB-D scene understanding algorithms. However, existing datasets still cover only a limited number of views or a restricted scale of spaces. In this paper, we introduce Matterport3D, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400 RGB-D images of 90 building-scale scenes. Annotations are provided with surface reconstructions, camera poses, and 2D and 3D semantic segmentations. The precise global alignment and comprehensive, diverse panoramic set of views over entire buildings enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.
Motivation & Objective
- Address the lack of large-scale, diverse RGB-D indoor datasets for training scene understanding models.
- Provide a globally aligned, building-scale RGB-D dataset with panoramic views and rich semantic annotations.
- Enable a range of learning tasks (keypoint matching, view overlap prediction, normal estimation, region classification, semantic voxel labeling) and establish baselines.
- Demonstrate how the dataset improves learning of descriptors, loop-closure, normals, and semantic understanding across tasks.
Proposed method
- Tripod-based Matterport capture yields 18 RGB-D images per panorama across 6 orientations with HDR color.
- Global bundle adjustment and textured mesh reconstruction provide 6-DoF camera poses and aligned surface representations.
- Crowdsourced and expert-verified 3D instance-level semantic annotations across 40 object categories.
- Baseline experiments demonstrating learning advantages for keypoint descriptors, view overlap prediction, surface normal estimation, region-type classification, and semantic voxel labeling.
Experimental results
Research questions
- RQ1Can Matterport3D pretrain and improve deep local descriptors for robust keypoint matching across diverse indoor views?
- RQ2Can comprehensive panoramic sampling enable effective loop-closure learning for view overlap prediction?
- RQ3Does training with high-quality Matterport3D depth improve surface normal estimation and generalize to other datasets?
- RQ4How does image field of view (single vs panorama) affect region-type classification performance?
- RQ5What is the performance of semantic voxel labeling on Matterport3D, and how does it compare to prior datasets?
Key findings
- Pretraining on Matterport3D yields improved keypoint matching performance on SUN3D benchmarks when using a ResNet-50 descriptor.
- View overlap prediction benefits from Matterport3D data, achieving higher retrieval metrics, with additional overlap regression loss providing further gains.
- Surface normal estimation improves when models are pretrained on Matterport3D and then evaluated on NYUv2, with MP pretraining showing better qualitative and quantitative results and cross-dataset generalization.
- Region-type classification benefits from panorama views, with increased field of view improving accuracy for several region categories (e.g., office, hallway, stairs, kitchen, etc.).
- Semantic voxel labeling on Matterport3D test scenes achieves an average accuracy of 70.3% across 20 classes, demonstrating strong 3D semantic understanding.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.