Skip to main content
QUICK REVIEW

[Paper Review] Towards Generating Virtual Movement from Textual Instructions A Case Study in Quality Assessment

Himangshu Sarma, Robert Porzel|arXiv (Cornell University)|Jun 6, 2020
Human Motion and Animation4 references4 citations
TL;DR

This study investigates human-based quality assessment of virtual movements generated from textual exercise instructions using crowdsourcing. It finds that RGB video visualization yields the highest inter-rater agreement (kappa = 0.74 for forward lunges), supporting its use as the optimal modality for reliable motion quality evaluation in automated pipeline development for exergames and therapy applications.

ABSTRACT

Many application areas ranging from serious games for health to learning by demonstration in robotics, could benefit from large body movement datasets extracted from textual instructions accompanied by images. The interpretation of instructions for the automatic generation of the corresponding motions (e.g. exercises) and the validation of these movements are difficult tasks. In this article we describe a first step towards achieving automated extraction. We have recorded five different exercises in random order with the help of seven amateur performers using a Kinect. During the recording, we found that the same exercise was interpreted differently by each human performer even though they were given identical textual instructions. We performed a quality assessment study based on that data using a crowdsourcing approach and tested the inter-rater agreement for different types of visualizations, where the RGBbased visualization showed the best agreement among the annotators.

Motivation & Objective

  • To explore human computation as a viable approach for assessing the quality of human body movements derived from textual exercise instructions.
  • To identify the most reliable visualization modality for human raters to evaluate motion quality in a crowdsourced setting.
  • To establish a foundational step toward automating the generation of virtual movements from exercise instruction sheets for use in therapy and gaming applications.
  • To assess inter-rater reliability across different motion visualization types (RGB, depth, skeleton, VR) in a quality assessment task.
  • To determine whether consistent quality judgments can be achieved despite variability in human interpretation of identical instructions.

Proposed method

  • Collected a Physical Exercise Instruction Sheet Corpus (PEISC) with ~1000 exercise descriptions categorized by body position (e.g., standing, sitting).
  • Recorded five equipment-free exercises using a Kinect from seven amateur performers (3 male, 4 female; M=25, SD=5), with ten iterations per performer.
  • Generated four visualization types from Kinect data: RGB (color video), depth maps, skeleton data, and virtual reality renderings.
  • Designed a crowdsourcing survey where 20 participants assessed motion quality by iteratively eliminating the worst and selecting the best performance across all visualizations.
  • Used Kappa statistics to measure inter-rater agreement across different visualizations and exercises.
  • Collected demographic data and post-exercise comprehension feedback to assess instruction clarity and performance variability.

Experimental results

Research questions

  • RQ1Can human computation reliably assess the quality of human body movements generated from textual exercise instructions?
  • RQ2Which motion visualization modality (RGB, depth, skeleton, VR) yields the highest inter-rater agreement in quality assessment tasks?
  • RQ3How consistent are human judgments of motion quality when evaluating the same exercise performed by different individuals with identical instructions?
  • RQ4To what extent does the clarity of textual instructions influence the variability in movement execution and quality assessment?
  • RQ5Can crowdsourced quality assessment serve as a valid intermediate step toward automating the pipeline from text to virtual motion generation?

Key findings

  • The RGB video visualization achieved the highest inter-rater agreement, with a Kappa statistic of 0.74 for the forward lunges exercise, indicating substantial agreement among annotators.
  • Inter-rater agreement varied across exercises, with the highest Kappa value (0.74) observed for forward lunges and the lowest (0.51) for squats, indicating moderate to substantial reliability.
  • Despite identical textual instructions, performers interpreted the same exercise differently, demonstrating significant inter-personal variability in movement execution.
  • Instruction comprehension was inconsistent—some participants found the sheets difficult to understand, while others found them easy, indicating variability in instruction clarity.
  • The study confirms that motion quality assessment is a challenging task even for humans, with moderate to substantial inter-rater reliability depending on the exercise and visualization type.
  • RGB video was perceived as the most intuitive and reliable modality for quality assessment, likely due to familiarity and naturalistic representation of motion.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.