[论文解读] HDHumans: A Hybrid Approach for High-fidelity Digital Humans
HDHumans 提出了一种混合框架,通过将一个形变的3D网格模板与神经辐射场(NeRF)紧密集成,实现了从多视角视频生成高保真、照片级真实感数字人像。通过相互引导——网格提升NeRF的采样效率和泛化能力,而NeRF则增强表面细节和监督——该方法实现了4K分辨率、时间上一致的新型视图与动作合成,即使在松垮衣物和未见动作下也能保持高质量表现。
Photo-real digital human avatars are of enormous importance in graphics, as they enable immersive communication over the globe, improve gaming and entertainment experiences, and can be particularly beneficial for AR and VR settings. However, current avatar generation approaches either fall short in high-fidelity novel view synthesis, generalization to novel motions, reproduction of loose clothing, or they cannot render characters at the high resolution offered by modern displays. To this end, we propose HDHumans, which is the first method for HD human character synthesis that jointly produces an accurate and temporally coherent 3D deforming surface and highly photo-realistic images of arbitrary novel views and of motions not seen at training time. At the technical core, our method tightly integrates a classical deforming character template with neural radiance fields (NeRF). Our method is carefully designed to achieve a synergy between classical surface deformation and NeRF. First, the template guides the NeRF, which allows synthesizing novel views of a highly dynamic and articulated character and even enables the synthesis of novel motions. Second, we also leverage the dense pointclouds resulting from NeRF to further improve the deforming surface via 3D-to-3D supervision. We outperform the state of the art quantitatively and qualitatively in terms of synthesis quality and resolution, as well as the quality of 3D surface reconstruction.
研究动机与目标
- 解决现有数字人像合成方法在处理高分辨率、新型视图、未见动作以及松垮衣物方面的局限性。
- 克服仅基于图像空间的方法存在的3D不一致与低保真度问题,以及仅使用网格或仅使用NeRF方法的泛化能力有限的问题。
- 通过一种紧密耦合的混合表示,实现从多视角视频高效、高保真地合成照片级真实感人像。
- 在4K分辨率下,实现3D表面重建与新型视图合成的最先进性能。
- 支持动作泛化与动态衣物形变,且无需为每种动作或服装类型提供显式监督。
提出的方法
- 提出一种混合表示方法,结合密集的、受动作驱动的3D形变网格模板与定义在表面薄壳区域内的神经辐射场(NeRF)。
- 利用网格作为几何先验,指导NeRF的采样与特征聚合,将每条光线的采样数从64+128减少至32,训练速度提升6倍。
- 利用NeRF的高频细节与显式3D监督,通过NeRF估计点的3D到3D监督,优化网格形变网络。
- 通过双向协同训练:网格引导NeRF实现高效、可泛化于动作的渲染;NeRF则提升网格细节并减少局部极小值。
- 以时空一致的方式参数化NeRF,基于骨骼运动与相机位姿进行条件控制,实现新型动作与视图的合成。
- 在训练过程中应用4K多视角监督,得益于网格引导的采样策略,实现了更高分辨率的输出。
实验结果
研究问题
- RQ1混合网格-NeRF表示能否共同提升高保真数字人像的3D表面重建与新型视图合成质量?
- RQ2形变网格模板如何用于引导NeRF采样,以提升效率并增强对未见动作与姿态的泛化能力?
- RQ3NeRF监督在多大程度上能提升学习到的形变网格的准确性与细节,特别是在松垮衣物与高频几何结构方面?
- RQ4该混合方法能否在保持时间一致性与动作泛化能力的同时,实现4K分辨率的照片级真实感渲染?
- RQ5此类紧密耦合系统在训练效率与推理速度方面的实际极限是什么?如何进一步优化?
主要发现
- HDHumans在3D表面重建与新型视图合成方面均达到最先进性能,显著提升了细节清晰度并减少了几何伪影。
- 该方法将NeRF每条光线的采样数从64+128减少至32,将4K多视角视频的训练时间从29天缩短至10天。
- 在4112×3008分辨率图像上的定量评估证实,4K监督可显著降低重建误差并提升保真度。
- 该方法成功合成了新型视图与未见动作的照片级真实感图像,包括如裙子等复杂动态衣物。
- 动作重定向结果表现出高度真实感,即使将蓝色连衣裙的动作迁移到黄色裙子上,也能完整保留褶皱等精细细节。
- 视频合成应用展示了该方法无缝将照片级真实感人像嵌入动态场景的能力,保持光照与几何的一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。