Skip to main content
QUICK REVIEW

[Paper Review] A Hybrid Framework for Action Recognition in Low-Quality Video Sequences

Tej Singh, Dinesh Kumar Vishwakarma|arXiv (Cornell University)|Mar 11, 2019
Human Pose and Action Recognition29 references5 citations
TL;DR

This paper proposes a hybrid framework for robust human action recognition in low-quality video sequences, combining sub-image histogram equalization for illumination enhancement and k-key pose silhouettes for feature extraction. The method achieves high recognition accuracy on degraded datasets like Weizmann, KTH, and Ballet Movement, outperforming existing techniques under low-exposure conditions while maintaining performance comparable to state-of-the-art methods.

ABSTRACT

Vision-based activity recognition is essential for security, monitoring and surveillance applications. Further, real-time analysis having low-quality video and contain less information about surrounding due to poor illumination, and occlusions. Therefore, it needs a more robust and integrated model for low quality and night security operations. In this context, we proposed a hybrid model for illumination invariant human activity recognition based on sub-image histogram equalization enhancement and k-key pose human silhouettes. This feature vector gives good average recognition accuracy on three low exposure video sequences subset of original actions video datasets. Finally, the performance of the proposed approach is tested over three manually downgraded low qualities Weizmann action, KTH, and Ballet Movement dataset. This model outperformed on low exposure videos over existing technique and achieved comparable classification accuracy to similar state-of-the-art methods.

Motivation & Objective

  • To address the challenge of human action recognition in low-quality video sequences with poor illumination and occlusions.
  • To develop a robust, illumination-invariant model for real-time surveillance and monitoring applications.
  • To enhance feature representation in low-exposure videos through image preprocessing and pose-based silhouettes.
  • To evaluate performance on manually downgraded versions of standard action recognition datasets.
  • To achieve competitive classification accuracy compared to state-of-the-art methods despite low input quality.

Proposed method

  • Applying sub-image histogram equalization to enhance contrast and visibility in low-exposure video frames.
  • Extracting k-key pose human silhouettes from enhanced frames to represent temporal action patterns.
  • Constructing a feature vector from the silhouettes for classification using a machine learning model.
  • Using a hybrid approach that combines image enhancement and pose-based representation to improve robustness.
  • Testing the framework on three manually downgraded datasets: Weizmann Action, KTH, and Ballet Movement.
  • Employing standard evaluation protocols to compare performance against existing methods.

Experimental results

Research questions

  • RQ1Can sub-image histogram equalization effectively improve recognition accuracy in low-exposure video sequences?
  • RQ2How does the use of k-key pose silhouettes enhance feature representation in low-quality videos?
  • RQ3Does the hybrid framework outperform existing methods on degraded video datasets?
  • RQ4To what extent does the proposed method maintain accuracy comparable to state-of-the-art techniques under low-quality conditions?
  • RQ5How robust is the framework to illumination variations and occlusions in surveillance scenarios?

Key findings

  • The proposed framework achieved higher average recognition accuracy than existing techniques on three low-exposure video sequence subsets.
  • The method outperformed baseline approaches on manually downgraded versions of the Weizmann Action, KTH, and Ballet Movement datasets.
  • Recognition accuracy remained comparable to state-of-the-art methods despite significant degradation in input video quality.
  • Sub-image histogram equalization effectively mitigated the impact of poor illumination on feature extraction.
  • The k-key pose silhouette representation provided stable and discriminative features under low-quality conditions.
  • The hybrid model demonstrated robustness in real-time surveillance applications with limited visual information.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.