Skip to main content
QUICK REVIEW

[논문 리뷰] Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification

Yang Du, Chunfeng Yuan|arXiv (Cornell University)|2018. 08. 03.
Human Pose and Action Recognition참고 문헌 53인용 수 43
한 줄 요약

본 논문은 다중 스케일 공간 피라미드를 갖춘 상호작용 인지 PCA-영감을 받은 자기 주의(attention)를 제안하고, 이를 비디오 동작 분류를 위한 시공간 주의로 확장하여, 인기 데이터셋에서 최첨단 성능을 달성한다.

ABSTRACT

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal elements of each local feature to obtain its weight score, which ignores interactions among local features. To address this, we propose an effective interaction-aware self-attention model inspired by PCA to learn attention maps. Furthermore, since different layers in a deep network capture feature maps of different scales, we use these feature maps to construct a spatial pyramid and then utilize multi-scale information to obtain more accurate attention scores, which are used to weight the local features in all spatial positions of feature maps to calculate attention maps. Moreover, our spatial pyramid attention is unrestricted to the number of its input feature maps so it is easily extended to a spatio-temporal version. Finally, our model is embedded in general CNNs to form end-to-end attention networks for action classification. Experimental results show that our method achieves the state-of-the-art results on the UCF101, HMDB51 and untrimmed Charades.

연구 동기 및 목표

  • CNN 특징 맵에서 이웃하는 로컬 특징들 간의 상호작용을 포착하여 행동 인식 성능 향상을 자극한다.

제안 방법

  • 특징 상호작용을 활용하기 위해 PCA에서 영감을 얻은 상호작용 인지 자기 주의 메커니즘을 도입한다.
  • 여러 CNN 계층으로부터 공간 특징 피라미드를 구성하여 다중 스케일 주의 맵을 얻는다.
  • 피라미드 특징에 걸쳐 채널 단위의 자기 주의를 수행하고 융합 함수를 사용하여 주의 맵을 계산한다.
  • 다중 비디오 프레임을 다루기 위해 공간 피라미드 주의를 시공간 버전으로 확장한다.
  • 주의를 정규화하고 CNN 내에서 엔드-투-엔드 학습이 가능하도록 상호작용 및 주의 손실을 도입한다.

실험 결과

연구 질문

  • RQ1로컬 CNN 특징들 간의 상호작용을 자기 주의에 어떻게 통합하여 행동 인식을 향상시킬 수 있는가?
  • RQ2다중 스케일 특징 맵에 대한 공간 피라미드가 비디오 프레임에 대해 더 정확한 주의 점수를 산출하는가?
  • RQ3주의 메커니즘이 임의 길이의 프레임 시퀀스에서 작동하는 시공간 형태로 확장될 수 있는가?
  • RQ4제안된 손실 함수와 다중 스케일 융합이 서로 다른 백본 네트워크에서 행동 분류를 개선하는가?

주요 결과

  • 이 방법은 UCF101, HMDB51, 그리고 잘리지 않은 Charades 데이터셋에서 최첨단 성능을 달성한다.
  • 3-스케일 공간 피라미드는 단일 스케일 주의에 비해 주목할 만한 성능 향상을 제공한다.
  • 융합 함수로 원소별 곱셈이 후보들 중에서 강한 성능을 보인다.
  • 이 방법은 CNN 백본(VGGNet-16, BN-Inception, Inception-ResNet-V2) 전반에서 일반성을 보여준다.
  • 시계열 확장은 K 프레임에 걸쳐 집계를 가능하게 하며, 테스트 시 더 많은 프레임을 사용할 때 향상이 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.