Skip to main content
QUICK REVIEW

[Paper Review] Dynamics-Regulated Kinematic Policy for Egocentric Pose Estimation

Zhengyi Luo, Ryo Hachiuma|arXiv (Cornell University)|Jun 10, 2021
Human Pose and Action Recognition48 references38 citations
TL;DR

This work combines kinematics and dynamics with scene context to estimate physically plausible 3D egocentric poses from a single head-mounted camera, using dynamics-regulated training to align a kinematic policy with a learned universal humanoid controller in a physics simulator.

ABSTRACT

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components are used disjointly, we synergize the two approaches via dynamics-regulated training. At each timestep, a kinematic model is used to provide a target pose using video evidence and simulation state. Then, a prelearned dynamics model attempts to mimic the kinematic pose in a physics simulator. By comparing the pose instructed by the kinematic model against the pose generated by the dynamics model, we can use their misalignment to further improve the kinematic model. By factoring in the 6DoF pose of objects (e.g., chairs, boxes) in the scene, we demonstrate for the first time, the ability to estimate physically-plausible 3D human-object interactions using a single wearable camera. We evaluate our egocentric pose estimation method in both controlled laboratory settings and real-world scenarios.

Motivation & Objective

  • Motivate and address the challenge of estimating physically plausible 3D egocentric full-body pose and object interaction from a single front-facing camera.
  • Develop a general-purpose humanoid physics controller learned from large MoCap data to mimic diverse human motions.
  • Propose a dynamics-regulated training procedure that tightly integrates kinematics, dynamics, and scene context for robust egocentric pose estimation.
  • Demonstrate how object context and scene constraints improve absolute 3D pose and interaction realism.
  • Evaluate on controlled MoCap lab data and real-world sequences to show improved pose and physics-based metrics.

Proposed method

  • Learn a Universal Humanoid Controller (UHC) from large MoCap data to mimic diverse human motions in a physics simulator.
  • Develop an object-aware kinematic policy that outputs per-frame target poses guided by video evidence and scene context.
  • Introduce dynamics-regulated training that combines supervised learning (kinematic targets) with reinforcement learning (physical plausibility via simulation).
  • Use an auto-regressive, agent-centric kinematic policy that fuses initial scene context, object state, and next-frame observations to predict per-step targets.
  • Integrate a pose-based loss for supervised learning and a physics-based reward that encourages motion imitation and dynamics consistency (via the UHC).
  • At test time, roll out the kinematic policy inside the physics simulator with the UHC acting as the low-level controller to produce physically-plausible poses.

Experimental results

Research questions

  • RQ1Can physically plausible 3D egocentric poses and human–object interactions be estimated from a single front-facing camera by integrating kinematics, dynamics, and scene context?
  • RQ2Does a learned, task-agnostic humanoid controller enable robust imitation of broad human motions in a physics simulator when guided by an object-aware kinematic policy?
  • RQ3Does dynamics-regulated training improve robustness to real-world domain shifts and improve absolute pose tracking compared to purely kinematic or purely dynamics-based methods?
  • RQ4How does incorporating 6DoF object poses and scene context affect the quality of egocentric pose estimation and interaction realism?

Key findings

  • The method estimates physically plausible 3D egocentric poses and interactions from a single head-mounted camera.
  • A Universal Humanoid Controller learned from MoCap can imitate a broad range of motions inside a physics simulator.
  • Dynamics-regulated training synergizes kinematics, dynamics, and scene context to improve robustness and pose accuracy.
  • Experiments on a MoCap dataset and real-world data show improvements over state-of-the-art methods on pose-based and physics-based metrics.
  • The approach explicitly models 6DoF object poses in the scene to enable realistic human–object interactions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.