[论文解读] HSPACE: Synthetic Parametric Humans Animated in Complex Environments
HSPACE 引入了一个大规模、逼真的合成数据集,包含多样化、参数化变化的人体在复杂室内外环境中的动画,通过 GHUM 身体模型实现真实的运动和完整的 3D 真值。当与弱监督的真实数据结合时,该数据集显著提升了 3D 人体姿态和形状估计的性能,尤其在模型容量增加时效果更明显。
Advances in the state of the art for 3d human sensing are currently limited by the lack of visual datasets with 3d ground truth, including multiple people, in motion, operating in real-world environments, with complex illumination or occlusion, and potentially observed by a moving camera. Sophisticated scene understanding would require estimating human pose and shape as well as gestures, towards representations that ultimately combine useful metric and behavioral signals with free-viewpoint photo-realistic visualisation capabilities. To sustain progress, we build a large-scale photo-realistic dataset, Human-SPACE (HSPACE), of animated humans placed in complex synthetic indoor and outdoor environments. We combine a hundred diverse individuals of varying ages, gender, proportions, and ethnicity, with hundreds of motions and scenes, as well as parametric variations in body shape (for a total of 1,600 different humans), in order to generate an initial dataset of over 1 million frames. Human animations are obtained by fitting an expressive human body model, GHUM, to single scans of people, followed by novel re-targeting and positioning procedures that support the realistic animation of dressed humans, statistical variation of body proportions, and jointly consistent scene placement of multiple moving people. Assets are generated automatically, at scale, and are compatible with existing real time rendering and game engines. The dataset with evaluation server will be made available for research. Our large-scale analysis of the impact of synthetic data, in connection with real data and weak supervision, underlines the considerable potential for continuing quality improvements and limiting the sim-to-real gap, in this practical setting, in connection with increased model capacity.
研究动机与目标
- 解决在复杂环境中缺乏大规模、多样化、逼真 3D 数据集的问题,这些数据集需包含多人运动且具备完整的 3D 真值。
- 通过创建可扩展、高保真的合成数据集并附带丰富标注,弥合 3D 人体姿态和形状估计中的模拟到真实(sim-to-real)差距。
- 支持开发和评估能够泛化于多样化体型、动作和场景条件的模型。
- 支持使用弱监督和合成数据进行模型训练,从而提升在真实世界基准上的性能。
- 提供一个可扩展、自动化的流水线,用于在复杂场景中生成多样化、无碰撞且视觉逼真的人体动画。
提出的方法
- 将 GHUM 3D 身体模型拟合到 100 个不同穿着个体的 3D 扫描数据上,实现参数化体型变化。
- 使用来自 CMU Mocap 的 100 个真实人体动作捕捉序列,对拟合后的 GHUM 模型进行重定向和动画处理。
- 自动将多个动画化的人体放置到 100 个合成的室内外场景中,场景包含逼真的光照和遮挡效果。
- 使用高保真游戏引擎渲染 4K/HDR 图像和视频,确保相机运动一致且视角多样化。
- 在场景布置过程中,自动实现人体之间以及人体与环境之间的碰撞避免。
- 生成丰富的标注信息,包括 3D 姿态/体型、人体分割、身体部位定位以及时间对应关系。
实验结果
研究问题
- RQ1大规模合成数据是否具备完整的 3D 真值,能够提升真实场景中 3D 人体姿态和形状估计的性能?
- RQ2当与弱监督的真实数据结合时,合成数据在减少模拟到真实域差距方面的有效性如何?
- RQ3在混合使用合成与真实数据进行训练时,增加模型容量在多大程度上能提升性能?
- RQ4体型多样性、动作多样性以及场景复杂性在多大程度上影响模型的泛化能力和鲁棒性?
- RQ5自动化、可扩展的流水线是否能够生成高保真、无碰撞的人体动画,并在复杂环境中提供准确的 3D 标注?
主要发现
- 在 HSPACE 上进行训练显著提升了 3D 人体姿态和形状估计性能,在 Human3.6M 基准测试中将 MPJPE-PA 降低至 39.0,超越了之前的 SOTA 水平。
- 在 HSPACE 数据上微调 THUNDR 模型后,MPJPE-PA 从 39.8 降低至 39.0,MPJPE-T 从 143.9 降低至 132.5(协议 P1 下)。
- 在 HITI(真实)和 HSPACE(合成)数据混合训练的模型中,性能得到提升,最大 T-THUNDR 模型的 MPJPE-PA 降低至 47。
- 当模型容量从 SMALL 提升至 BIG 时,在 HSPACE 测试集上的 MPJPE-PA 从 50 降低至 47,证明了更大架构的优势。
- 包含完整 3D 监督的合成数据(S#3D)带来了显著的性能提升,尤其在与真实数据和弱监督结合时效果更明显。
- 该数据集在 HSPACE-TEST 和 HITI-TEST 测试集上均表现出色,表明其在评估模型泛化能力方面具有重要价值,超越了标准基准。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。