[论文解读] Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
HM3D 是一个大规模、高保真度的 3D 室内数据集,包含 1000 个建筑规模的重建,相比早期数据集在规模、完整性和视觉保真度方面具有更高的水平,从而在多个测试环境中提升具象化 AI 导航能力。
We present the Habitat-Matterport 3D (HM3D) dataset. HM3D is a large-scale dataset of 1,000 building-scale 3D reconstructions from a diverse set of real-world locations. Each scene in the dataset consists of a textured 3D mesh reconstruction of interiors such as multi-floor residences, stores, and other private indoor spaces. HM3D surpasses existing datasets available for academic research in terms of physical scale, completeness of the reconstruction, and visual fidelity. HM3D contains 112.5k m^2 of navigable space, which is 1.4 - 3.7x larger than other building-scale datasets such as MP3D and Gibson. When compared to existing photorealistic 3D datasets such as Replica, MP3D, Gibson, and ScanNet, images rendered from HM3D have 20 - 85% higher visual fidelity w.r.t. counterpart images captured with real cameras, and HM3D meshes have 34 - 91% fewer artifacts due to incomplete surface reconstruction. The increased scale, fidelity, and diversity of HM3D directly impacts the performance of embodied AI agents trained using it. In fact, we find that HM3D is `pareto optimal' in the following sense -- agents trained to perform PointGoal navigation on HM3D achieve the highest performance regardless of whether they are evaluated on HM3D, Gibson, or MP3D. No similar claim can be made about training on other datasets. HM3D-trained PointNav agents achieve 100% performance on Gibson-test dataset, suggesting that it might be time to retire that episode dataset.
研究动机与目标
- 提供一个大规模、近完整的真实世界室内建筑的 3D 重建数据集。
- 提高视觉保真度并降低相对于现有数据集的重建伪影。
- 展示 HM3D 在导航任务中用于训练和评估具身 AI 代理的效用。
- 展示在多个测试环境中 HM3D 训练的代理的帕累托最优迁移性能。
提出的方法
- 使用 Matterport Pro2 扫描和 Matterport 流程捕获并重建 1000 个建筑规模的室内空间。
- 通过定量指标量化相对于先前数据集的规模、完整性和视觉保真度。
- 评估在 HM3D 与其他数据集相比训练时具身 AI 导航(PointGoal/PointNav)的性能。
- 提供可用于 Habitat 的元数据和集成,便于在 Habitat 仿真器中训练。
实验结果
研究问题
- RQ1与现有室内数据集相比,HM3D 是否在训练具身 AI 代理方面提供可扩展的优势?
- RQ2HM3D 的规模、完整性和保真度如何影响代理的导航性能以及跨数据集的泛化?
- RQ3在 Gibson、MP3D 与 HM3D 测试集上的性能是否在帕累托意义上最优?
- RQ4HM3D 的视觉保真度和重建完整度与真实世界影像及早前数据集相比如何?
- RQ5在训练场景中增加可导航区域是否与 PointNav 性能的提升相关?
主要发现
- HM3D 提供 1000 个建筑尺度的重建,具备 112.5k m^2 的可导航空间,在规模上超过 MP3D 与 Gibson。
- HM3D 的重建伪影减少了 34.91%,视觉保真度提高了 20.85%。
- 经 HM3D 训练的 PointNav 代理在 Gibson、MP3D 和 HM3D 测试集上达到最先进或帕累托最优的性能,包括 Gibson 测试在深度输入下的 100% 成功。
- 渲染的 HM3D 图像相对于 MP3D 和 Gibson 的真实图像基线在 FID/KID 分数显著更低,表明对真实影像具有更高的视觉保真度。
- 导航性能与训练 HM3D 场景的总可导航面积近线性相关(rho≈0.88)。
- HM3D 代理对更难的情节具有更好的泛化能力,展示了多样化的布局和外观有助于迁移性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。