Skip to main content
QUICK REVIEW

[论文解读] StereoPIFu: Depth Aware Clothed Human Digitization via Stereo Vision

H. C. Yang, Juyong Zhang|arXiv (Cornell University)|Apr 12, 2021
Advanced Vision and Imaging参考文献 55被引用 4
一句话总结

StereoPIfu 提出了一种基于双目视觉和隐式函数表示的深度感知三维穿衣人体重建方法。通过整合双目网络的体素对齐特征与相对 z 偏移机制,其在精度、完整性和鲁棒性方面均优于先前的单图像方法(如 PIFuHD),将点到表面误差降低 74.1%,Chamfer 距离降低 75.1%。

ABSTRACT

In this paper, we propose StereoPIFu, which integrates the geometric constraints of stereo vision with implicit function representation of PIFu, to recover the 3D shape of the clothed human from a pair of low-cost rectified images. First, we introduce the effective voxel-aligned features from a stereo vision-based network to enable depth-aware reconstruction. Moreover, the novel relative z-offset is employed to associate predicted high-fidelity human depth and occupancy inference, which helps restore fine-level surface details. Second, a network structure that fully utilizes the geometry information from the stereo images is designed to improve the human body reconstruction quality. Consequently, our StereoPIFu can naturally infer the human body's spatial location in camera space and maintain the correct relative position of different parts of the human body, which enables our method to capture human performance. Compared with previous works, our StereoPIFu significantly improves the robustness, completeness, and accuracy of the clothed human reconstruction, which is demonstrated by extensive experimental results.

研究动机与目标

  • 解决单图像三维人体重建方法(如 PIFu)中存在的深度模糊与空间定位不一致问题。
  • 提升穿衣人体重建中的几何细节恢复与结构一致性。
  • 利用双目视觉提供比单视角方法更丰富的几何约束。
  • 通过深度感知监督消除背侧区域的伪影(如肢体断裂与几何复制现象)。
  • 实现在无需规范空间归一化的情况下,对复杂人体形态与姿态进行精确、鲁棒且完整的三维重建。

提出的方法

  • 引入从基于双目的体素数据中提取的体素对齐特征,实现深度感知的三维特征表示。
  • 从校正后的双目图像特征构建代价体,以编码对应像素之间的几何相关性。
  • 在体素网格上使用三线性插值,为三维查询点生成空间对齐的特征。
  • 将三维查询点与其投影像素预测深度之间的相对 z 偏移作为隐式函数的输入。
  • 利用预测的高保真深度图作为全局形状先验,引导占据推理并提升表面细节。
  • 设计一种充分挖掘双目几何信息的网络架构,实现在相机空间中的空间定位,而无需规范归一化。

实验结果

研究问题

  • RQ1与单图像方法相比,双目视觉是否能提升三维穿衣人体重建中的深度精度与空间一致性?
  • RQ2来自双目网络的体素对齐特征在基于隐式函数的重建中,如何增强几何细节恢复能力?
  • RQ3与绝对 z 值相比,使用相对 z 偏移在多大程度上能提升表面细节并减少伪影?
  • RQ4来自双目估计的深度感知监督是否能消除重建人体中如肢体断裂等结构错误?
  • RQ5与最先进方法相比,该方法在复杂姿态、遮挡情况及真实世界数据上的表现如何?

主要发现

  • StereoPIfu 将平均点到表面距离从 PIFuHD 的 1.97cm 降低至 0.51cm,提升 74.1%。
  • Chamfer 距离从 PIFuHD 的 2.21cm 降低至 0.55cm,减少 75.1%,表明几何精度显著更优。
  • 该方法成功重建了如孕妇和弯曲姿势等复杂形态,准确保持了肢体间的相对位置关系。
  • 在使用双目相机捕获的真实世界数据上,StereoPIfu 即使在光照与标定条件变化下,也能生成稳定且精确的重建结果。
  • 与 PIFuHD 相比,StereoPIfu 避免了如腿部过长与肢体比例错误等结构伪影,尤其在非输入视角下更为明显。
  • 该方法在无需刚性对齐或缩放操作的情况下实现深度感知重建,而单图像基线方法则通常需要此类处理。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。