Skip to main content
QUICK REVIEW

[Paper Review] Hand tracking for clinical applications: validation of the Google MediaPipe Hand (GMH) and the depth-enhanced GMH-D frameworks

Gianluca Amprimo, Giulia Masi|arXiv (Cornell University)|Aug 2, 2023
Stroke Rehabilitation and RecoveryMedicine3 citations
TL;DR

This study validates Google MediaPipe Hand (GMH) and its depth-enhanced variant, GMH-D, for clinical hand tracking using an RGB-depth camera. GMH-D achieves superior spatial accuracy over GMH in measuring 3D hand movements during dynamic tasks, demonstrating strong temporal and spectral consistency with a gold-standard motion capture system, making it a reliable tool for clinical assessment of hand function.

ABSTRACT

Accurate 3D tracking of hand and fingers movements poses significant challenges in computer vision. The potential applications span across multiple domains, including human-computer interaction, virtual reality, industry, and medicine. While gesture recognition has achieved remarkable accuracy, quantifying fine movements remains a hurdle, particularly in clinical applications where the assessment of hand dysfunctions and rehabilitation training outcomes necessitate precise measurements. Several novel and lightweight frameworks based on Deep Learning have emerged to address this issue; however, their performance in accurately and reliably measuring fingers movements requires validation against well-established gold standard systems. In this paper, the aim is to validate the handtracking framework implemented by Google MediaPipe Hand (GMH) and an innovative enhanced version, GMH-D, that exploits the depth estimation of an RGB-Depth camera to achieve more accurate tracking of 3D movements. Three dynamic exercises commonly administered by clinicians to assess hand dysfunctions, namely Hand Opening-Closing, Single Finger Tapping and Multiple Finger Tapping are considered. Results demonstrate high temporal and spectral consistency of both frameworks with the gold standard. However, the enhanced GMH-D framework exhibits superior accuracy in spatial measurements compared to the baseline GMH, for both slow and fast movements. Overall, our study contributes to the advancement of hand tracking technology, the establishment of a validation procedure as a good-practice to prove efficacy of deep-learning-based hand-tracking, and proves the effectiveness of GMH-D as a reliable framework for assessing 3D hand movements in clinical applications.

Motivation & Objective

  • To validate the accuracy of the Google MediaPipe Hand (GMH) framework for 3D hand tracking in clinical settings.
  • To evaluate the performance of an enhanced version, GMH-D, which integrates depth estimation from an RGB-depth camera to improve spatial accuracy.
  • To assess both frameworks against a gold-standard motion capture system during clinically relevant hand movement tasks.
  • To establish a validation procedure as a best practice for deep learning-based hand tracking in clinical applications.

Proposed method

  • The study employs the GMH framework, a lightweight deep learning model for 2D hand keypoint detection, extended into GMH-D by fusing depth data from an RGB-depth camera.
  • The GMH-D framework uses a multi-stream architecture to combine RGB image features with depth map features for improved 3D keypoint regression.
  • Three dynamic hand exercises—Hand Opening-Closing, Single Finger Tapping, and Multiple Finger Tapping—were performed by participants and recorded using both the GMH/GMH-D systems and a gold-standard motion capture system.
  • Temporal and spectral consistency between the systems was evaluated using cross-correlation and spectral coherence analysis.
  • Spatial accuracy was quantified via root mean square error (RMSE) between predicted and ground-truth 3D keypoint positions.
  • The validation process included repeated trials under varying movement speeds to test robustness across slow and fast motions.

Experimental results

Research questions

  • RQ1How accurately does the GMH framework track 3D hand movements compared to a gold-standard motion capture system in clinical tasks?
  • RQ2To what extent does integrating depth data in GMH-D improve spatial accuracy over the baseline GMH framework?
  • RQ3How consistent are the temporal and spectral characteristics of GMH and GMH-D outputs relative to the gold standard?
  • RQ4Does the performance of GMH-D remain stable across both slow and fast hand movement velocities?

Key findings

  • GMH-D demonstrated significantly improved spatial accuracy compared to GMH, with lower root mean square error (RMSE) in 3D keypoint estimation across all tested movements.
  • Both GMH and GMH-D showed high temporal consistency, with cross-correlation coefficients exceeding 0.95 for all tasks.
  • Spectral coherence analysis revealed that both frameworks preserved the frequency content of hand movements with high fidelity, especially in the range relevant to clinical assessment.
  • GMH-D outperformed GMH in spatial accuracy for both slow and fast movements, indicating robustness across motion velocities.
  • The study establishes a reproducible validation pipeline for deep learning-based hand tracking systems in clinical environments.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.