[Paper Review] Human Action Recognition and Prediction: A Survey
This survey reviews state-of-the-art techniques for vision-based human action recognition and prediction from videos, covering representations, classifiers, datasets, challenges, applications, and future directions.
Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-based action recognition and prediction from videos are such tasks, where action recognition is to infer human actions (present state) based upon complete action executions, and action prediction to predict human actions (future state) based upon incomplete action executions. These two tasks have become particularly prevalent topics recently because of their explosively emerging real-world applications, such as visual surveillance, autonomous driving vehicle, entertainment, and video retrieval, etc. Many attempts have been devoted in the last a few decades in order to build a robust and effective framework for action recognition and prediction. In this paper, we survey the complete state-of-the-art techniques in action recognition and prediction. Existing models, popular algorithms, technical difficulties, popular action databases, evaluation protocols, and promising future directions are also provided with systematic discussions.
Motivation & Objective
- Survey the complete state-of-the-art techniques in action recognition and prediction.
- Clarify the distinctions and similarities between recognition (present state) and prediction (future state).
- Discuss representative datasets, evaluation protocols, and real-world applications.
- Identify key challenges and outline promising directions for future research.
Proposed method
- Organize action representation approaches into holistic and local features, and discuss their strengths and weaknesses.
- Review shallow (hand-crafted) and deep learning-based representations and classifiers.
- Categorize action classifiers into direct, sequential, space-time, and part-based models, including end-to-end deep frameworks.
- Summarize datasets, evaluation protocols, and practical challenges such as intra-/inter-class variation and data labeling.
- Highlight real-world applications and future research directions.
Experimental results
Research questions
- RQ1What are effective action representations for robust recognition and prediction across varying viewpoints and backgrounds?
- RQ2How do recognition and prediction tasks differ in terms of data requirements and temporal modeling?
- RQ3What are the main challenges hindering real-world deployment (variability, noise, data annotation) and how can they be mitigated?
- RQ4What directions hold the most promise for advancing action prediction in urgent scenarios (e.g., autonomous driving)?
Key findings
- The survey consolidates state-of-the-art techniques in action recognition and prediction and discusses datasets, protocols, and challenges.
- It highlights the distinction between recognizing complete actions and predicting actions from incomplete sequences.
- It surveys broad families of representations and classifiers, from holistic, local, trajectory-based, to part-based and deep learning approaches.
- It discusses real-world applications such as surveillance, video retrieval, entertainment, human-robot interaction, and autonomous driving.
- It identifies major challenges including intra-/inter-class variation, clutter, camera motion, data scarcity, action vocabulary, and uneven predictability.
- It outlines promising future directions and gaps in current research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.