Skip to main content
QUICK REVIEW

[论文解读] EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras

Helge Rhodin, Christian Richardt|Max Planck Digital Library|Sep 23, 2016
Human Pose and Action Recognition参考文献 68被引用 7
一句话总结

EgoCap 提出了一种使用两个头戴式鱼眼镜头的实时、无标记全身动作捕捉方法,结合生成式姿态估计模型与基于卷积神经网络(ConvNet)的身体部位检测器,实现在室内和室外环境中(包括人群密集场景)的非受限、第一视角动作捕捉,适用于沉浸式虚拟现实应用。

ABSTRACT

Marker-based and marker-less optical skeletal motion-capture methods use an outside-in arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort by possibly needed marker suits, and their recording volume is severely restricted and often constrained to indoor scenes with controlled backgrounds. Alternative suit-based systems use several inertial measurement units or an exoskeleton to capture motion. This makes capturing independent of a confined volume, but requires substantial, often constraining, and hard to set up body instrumentation. We therefore propose a new method for real-time, marker-less and egocentric motion capture which estimates the full-body skeleton pose from a lightweight stereo pair of fisheye cameras that are attached to a helmet or virtual reality headset. It combines the strength of a new generative pose estimation framework for fisheye views with a ConvNet-based body-part detector trained on a large new dataset. Our inside-in method captures full-body motion in general indoor and outdoor scenes, and also crowded scenes with many people in close vicinity. The captured user can freely move around, which enables reconstruction of larger-scale activities and is particularly useful in virtual reality to freely roam and interact, while seeing the fully motion-captured virtual body.

研究动机与目标

  • 解决传统外部定位动作捕捉系统存在的局限性,这些系统需要受控的拍摄空间和背景环境。
  • 克服可穿戴传感器系统(如惯性测量单元 IMU、外骨骼)的限制,后者需要大量设备部署和校准。
  • 利用轻量化、可穿戴的摄像头配置,在非受限、大规模且人群密集的环境中实现全身动作捕捉。
  • 开发一种方法,通过重建用户自身在三维空间中的身体形态,支持沉浸式虚拟现实中的具身交互。
  • 创建一个新的数据集和训练框架,用于第一视角鱼眼姿态估计,以支持在真实世界中稳定部署。

提出的方法

  • 在头盔或 VR 头显上部署一对鱼眼镜头的立体视觉系统,从第一视角捕捉全身图像。
  • 采用基于射线投射成像模型的生成式姿态估计框架,实现可微分的对齐与可见性建模。
  • 集成在大规模真实第一视角鱼眼图像数据集上自动标注后训练的基于卷积神经网络(ConvNet)的身体部位检测器。
  • 通过在场景背景上应用结构光流法(structure-from-motion)来推断全局运动,从而将局部身体姿态估计与全局头戴设备位置解耦。
  • 利用鱼眼镜头强烈的透视畸变增强深度和距离感知线索,提升跟踪鲁棒性。
  • 结合生成模型与 CNN 检测结果,对姿态和可见性进行联合优化,以处理自遮挡和复杂姿态情况。

实验结果

研究问题

  • RQ1轻量化、头戴式的双鱼眼镜头装置是否能够在非受限的室内外环境中实现全身无标记动作捕捉?
  • RQ2在高度畸变的鱼眼图像视图中,基于可微分可见性与对齐能量的生成式姿态估计模型效果如何?
  • RQ3在自遮挡和背景杂乱条件下,基于 CNN 的身体部位检测器若在合成与真实的第一视角鱼眼数据上进行训练,其姿态估计精度能提升多少?
  • RQ4在无外部传感器的情况下,能否通过场景背景的结构光流法可靠估计头戴设备的全局运动?
  • RQ5与传统的外部定位系统或 IMU 系统相比,第一视角鱼眼摄像头系统在设置自由度、用户舒适度和动态场景中跟踪鲁棒性方面表现如何?

主要发现

  • EgoCap 成功实现了仅使用两个头戴式鱼眼镜头的实时、无标记全身动作捕捉,在大规模和人群密集场景中表现出稳健性能。
  • 生成式姿态模型与基于 CNN 的检测器相结合,显著提升了跟踪精度,尤其在自遮挡和复杂身体构型下表现更优。
  • 通过在场景背景特征上应用结构光流法,实现了稳定的全局姿态估计,成功将局部身体姿态与全局运动解耦。
  • 由于鱼眼镜头强烈的透视畸变,提供了更优的深度感知线索,提升了对相机前后运动的检测能力。
  • 尽管绝对精度低于商业化的外部定位系统,EgoCap 展现出在非受限真实世界环境中动作捕捉的可行性与可扩展性。
  • 该方法通过从用户自身视角实现实时、全身的虚拟化身重建,支持沉浸式虚拟现实应用,增强了具身感与交互体验。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。