Skip to main content
QUICK REVIEW

[논문 리뷰] Detecting Violence in Video Based on Deep Features Fusion Technique

Heyam M. Bin Jahlan, Lamiaa A. Elrefaei|arXiv (Cornell University)|2022. 04. 15.
Anomaly Detection Techniques and Applications인용 수 7
한 줄 요약

이 논문은 비디오에서의 자동 폭력 탐지에 대해 두 개의 서로 다른 CNN인 AlexNet과 SqueezeNet의 특징을 융합하는 딥러닝 방법을 제안한다. 각각의 CNN 뒤에 ConvLSTM를 적용하여 시공간적 특징을 추출하고, 융합된 특징은 풀링된 후 완전 연결층을 통해 분류되어, 각각 Hockey Fight, Movie, Violent Flow 데이터셋에서 97%, 100%, 96%의 정확도를 기록하며 최신 기술을 초월한다.

ABSTRACT

With the rapid growth of surveillance cameras in many public places to mon-itor human activities such as in malls, streets, schools and, prisons, there is a strong demand for such systems to detect violence events automatically. Au-tomatic analysis of video to detect violence is significant for law enforce-ment. Moreover, it helps to avoid any social, economic and environmental damages. Mostly, all systems today require manual human supervisors to de-tect violence scenes in the video which is inefficient and inaccurate. in this work, we interest in physical violence that involved two persons or more. This work proposed a novel method to detect violence using a fusion tech-nique of two significantly different convolutional neural networks (CNNs) which are AlexNet and SqueezeNet networks. Each network followed by separate Convolution Long Short Term memory (ConvLSTM) to extract ro-bust and richer features from a video in the final hidden state. Then, making a fusion of these two obtained states and fed to the max-pooling layer. Final-ly, features were classified using a series of fully connected layers and soft-max classifier. The performance of the proposed method is evaluated using three standard benchmark datasets in terms of detection accuracy: Hockey Fight dataset, Movie dataset and Violent Flow dataset. The results show an accuracy of 97%, 100%, and 96% respectively. A comparison of the results with the state of the art techniques revealed the promising capability of the proposed method in recognizing violent videos.

연구 동기 및 목표

  • 감시 영상에서 수동 폭력 탐지의 비효율성과 부정확성을 해결하기 위해.
  • 두 명 이상의 개인이 관여하는 신체 폭력 행동을 탐지할 수 있는 자동화된 시스템을 개발하기 위해.
  • 두 개의 서로 다른 합성곱 신경망에서 유래한 보완적인 특징을 융합하여 탐지 정확도를 향상시키기 위해.
  • 표준 기준 데이터셋을 통해 성능을 철저히 평가하여 시스템의 견고성을 확보하기 위해.

제안 방법

  • 비디오 프레임에서 다양한 시각적 패턴을 캡처하기 위해 AlexNet과 SqueezeNet을 별도의 특징 추출기로 활용하기 위해.
  • 비디오 시퀀스에서 시간적 동역학을 모델링하고 강력한 시공간적 특징을 추출하기 위해 각 CNN 뒤에 ConvLSTM 네트워크를 적용하기 위해.
  • 두 개의 ConvLSTM 네트워크에서 유도된 최종 은닉 상태를 융합하여 상보적인 표현을 통합하기 위해.
  • 차원을 줄이고 분류 능력을 향상시키기 위해 융합된 특징 표현에 최대 풀링층을 적용하기 위해.
  • 최종 폭력 예측을 위해 풀링된 특징을 완전 연결층을 거쳐 소프트맥스 분류기로 분류하기 위해.

실험 결과

연구 질문

  • RQ1두 개의 서로 다른 CNN 아키텍처 융합이 비디오 시퀀스에서의 폭력 탐지 정확도 향상에 기여하는가?
  • RQ2다양한 CNN에 ConvLSTM를 통합하는 것이 폭력 행동의 시공간적 동역학을 얼마나 효과적으로 포착하는가?
  • RQ3제안된 특징 융합 전략이 표준 폭력 탐지 기준 데이터셋에서 기존 기술을 능가하는가?
  • RQ4이 하이브리드 딥러닝 아키텍처로 실세계 비디오 데이터셋에서 어떤 정도의 정확도를 달성할 수 있는가?

주요 결과

  • 제안된 방법은 Hockey Fight 데이터셋에서 97%의 탐지 정확도를 기록하여 스포츠 관련 폭력 사건에 대해 뛰어난 성능을 보였다.
  • Movie 데이터셋에서는 100%의 정확도를 달성하여 정제된 영화 폭력 영상에 대해 뛰어난 일반화 능력을 보였다.
  • Violent Flow 데이터셋에서는 96%의 정확도를 기록하여 다양한 비디오 콘텐츠에 걸쳐 뛰어난 견고성을 확인하였다.
  • AlexNet과 SqueezeNet의 특징를 ConvLSTM와 융합한 전략이 폭력 인식 작업에서 최신 기술을 능가하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.