[Paper Review] Intention from Motion.
This paper introduces Intention from Motion, a novel action prediction paradigm that infers human intentions solely from kinematic motion data—without contextual cues—by leveraging a new multi-modal dataset of 3D motion capture and 2D video. The method achieves reliable intention prediction using only motion kinematics, demonstrating that state-of-the-art action recognition models fail on this task and proposing a new classification technique tailored to intention inference from motion alone.
In this paper, we propose Intention from Motion, a new paradigm for action prediction where, without using any contextual information, we can predict human intentions all originating from the same motor act, non specific of the following performed action. To investigate such novel problem, we designed a proof of concept consisting in a new multi-modal dataset of motion capture marker 3D data and 2D video sequences where, by only analysing very similar movements in both training and test phases, we are able to predict the underlying intention, i.e., the future, never observed, action. Through an extended baseline assessment, we evaluate the proposed dataset employing state-of-the-art 3D and 2D action recognition techniques, while using fusion methods to fully exploit its multi-modal nature. We also report comparative benchmarking tests using existing action prediction pipelines, showing that such algorithms can not deal well with the proposed intention prediction problem. In the end, we demonstrate that intentions can be predicted in a reliable way, ultimately devising a novel classification technique to infer the intention from kinematic information only.
Motivation & Objective
- To investigate whether human intentions can be predicted from motion alone, without contextual or environmental information.
- To address the challenge of predicting diverse future actions from identical motor acts, a problem not well handled by existing action prediction pipelines.
- To design and release a new multi-modal dataset of 3D motion capture and 2D video sequences to support research in motion-only intention prediction.
- To evaluate the performance of state-of-the-art 3D and 2D action recognition models on this new task, revealing their limitations in handling intention prediction.
- To develop and validate a novel classification technique that infers intention exclusively from kinematic motion features.
Proposed method
- The authors created a multi-modal dataset combining 3D motion capture marker data and 2D video sequences, capturing very similar movements across different intentions.
- They applied state-of-the-art 3D and 2D action recognition models to the dataset, using early and late fusion strategies to exploit both modalities.
- The method evaluates performance on predicting the underlying intention from motion sequences that are visually and kinematically similar but lead to different future actions.
- A novel classification technique was devised to infer intention directly from kinematic features, bypassing the need for contextual or environmental context.
- Baseline models were tested on the dataset to assess their ability to generalize to intention prediction, revealing their shortcomings in this specific setting.
- The approach focuses on identifying subtle kinematic differences in motion that signal divergent intentions, even when the initial motion is nearly identical.
Experimental results
Research questions
- RQ1Can human intentions be reliably predicted from motion data alone, without any contextual or environmental information?
- RQ2How do existing action recognition and prediction pipelines perform when applied to the task of predicting intentions from identical motor acts?
- RQ3What are the key kinematic cues in motion that distinguish different future actions despite similar initial movements?
- RQ4To what extent can multi-modal fusion of 3D motion and 2D video improve intention prediction performance in this setting?
- RQ5Can a novel classification model trained solely on motion kinematics outperform existing methods on this intention prediction task?
Key findings
- Existing action recognition and prediction pipelines fail to generalize effectively to the task of predicting intentions from identical motor acts.
- The proposed multi-modal dataset enables reliable intention prediction when using motion-only features, demonstrating the feasibility of the Intention from Motion paradigm.
- State-of-the-art 3D and 2D action recognition models show limited performance on this novel task, indicating a gap in current approaches.
- Fusion of 3D motion capture and 2D video data improves performance, but the best results are achieved when using a dedicated classification model trained on kinematic features alone.
- A novel classification technique was successfully developed that infers intention exclusively from kinematic information, outperforming general-purpose models on this specific task.
- The results confirm that subtle differences in motion kinematics carry sufficient information to predict divergent future actions, even when initial movements are nearly identical.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.