Skip to main content
QUICK REVIEW

[논문 리뷰] PAN: Towards Fast Action Recognition via Learning Persistence of Appearance

Can Zhang, Yuexian Zou|arXiv (Cornell University)|2020. 08. 08.
Human Pose and Action Recognition참고 문헌 49인용 수 32
한 줄 요약

PAN은 광학 흐름을 대체하는 가볍고 지속성 있는 Appearance의 PA 큐와 다양한 시간척도 축적 풀링(VAP)을 활용하여 동작 인식을 위한 빠르고 엔드-투-엔드 프레임워크를 제시합니다. 이를 통해 장기 역동성을 모델링하고 실시간 가능성과 높은 정확도를 달성합니다.

ABSTRACT

Efficiently modeling dynamic motion information in videos is crucial for action recognition task. Most state-of-the-art methods heavily rely on dense optical flow as motion representation. Although combining optical flow with RGB frames as input can achieve excellent recognition performance, the optical flow extraction is very time-consuming. This undoubtably will count against real-time action recognition. In this paper, we shed light on fast action recognition by lifting the reliance on optical flow. Our motivation lies in the observation that small displacements of motion boundaries are the most critical ingredients for distinguishing actions, so we design a novel motion cue called Persistence of Appearance (PA). In contrast to optical flow, our PA focuses more on distilling the motion information at boundaries. Also, it is more efficient by only accumulating pixel-wise differences in feature space, instead of using exhaustive patch-wise search of all the possible motion vectors. Our PA is over 1000x faster (8196fps vs. 8fps) than conventional optical flow in terms of motion modeling speed. To further aggregate the short-term dynamics in PA to long-term dynamics, we also devise a global temporal fusion strategy called Various-timescale Aggregation Pooling (VAP) that can adaptively model long-range temporal relationships across various timescales. We finally incorporate the proposed PA and VAP to form a unified framework called Persistent Appearance Network (PAN) with strong temporal modeling ability. Extensive experiments on six challenging action recognition benchmarks verify that our PAN outperforms recent state-of-the-art methods at low FLOPs. Codes and models are available at: https://github.com/zhang-can/PAN-PyTorch.

연구 동기 및 목표

  • 시간이 많이 걸리는 광학 흐름에 대한 의존성을 줄여 빠른 동작 인식을 촉진한다.
  • 동작 경계에 초점을 맞춘 효율적인 모션 큐로서 Persistence of Appearance(PA)를 도입한다.
  • 가벼운 PA 모듈을 개발하고 이를 시간 융합 전략(VAP)과 결합해 PAN을 형성한다.
  • PA의 두 가지 인코딩 스킴을 탐색하고 성능과 효율성을 비교한다.
  • 시간 의존적인 데이터셋과 장면 의존적인 데이터셋에서 PAN의 효과를 입증한다.

제안 방법

  • PA를 이미지 공간의 밝기일관성을 특징 공간으로 확장하여 경계 움직임을 포착하기 위해 픽셀 단위의 특징 차이를 사용하여 정의한다.
  • 인접 프레임으로부터 PA 주목 맵을 생성하기 위해 단일 저수준 컨볼루션 계층(8×7×7 필터)을 갖춘 PA 모듈을 구현한다.
  • 두 가지 인코딩 스킴을 사용한다: e1(동작 양상으로서의 PA)과 e2(주목으로서의 PA); e1은 PA를 백본에 직접 입력하고, e2는 PA를 사용해 appearance 특징을 조절한다.
  • PAN은 두 가지 변형으로 구성한다: PAN Full(나중 융합이 있는 별도 RGB 및 PA 분기)과 PAN Lite(연결된 RGB+PA 입력을 처리하는 통합 백본).
  • 다중 시간 스케일 풀링을 통한 장기 시계열 정보를 융합하고 시간대 간 적응 가중치를 제공하는 Various-timescale Aggregation Pooling(VAP)을 도입한다.

실험 결과

연구 질문

  • RQ1PA가 광학 흐름보다 더 효율적으로 동작 인식에 필요한 동일한 모션 정보를 포착할 수 있는가?
  • RQ2깊은 백본과 글로벌 시간 융합 전략(VAP)과 통합될 때 PA가 시간 모델링에 어떻게 기여하는가?
  • RQ3어떤 PA 인코딩 스킴이 PAN에서 더 나은 인식 성능과 효율을 제공하는가?
  • RQ4PAN Full과 PAN Lite가 시간 의존형 데이터셋과 장면 의존형 데이터셋에서 보완적이거나 트레이드오프적인 성능을 제공하는가?

주요 결과

  • PA는 매우 효율적(8196 fps)이며 UCF101 분할 1에서 강한 정확도(89.5%)를 달성하여 런타임 측면에서 다수의 광학 흐름 기반 접근법을 능가하고 경쟁력 있는 정확도를 유지한다.
  • PA를 모션 양상으로 인코딩(e1)하는 것이 주의로 인코딩하는(e2) 것보다 UCF101에서 정확도와 효율성 모두에서 우수하다; e2는 느리고 정확도도 약간 낮다.
  • PAN Full과 PAN Lite는 시간 의존적 데이터셋과 장면 의존적 데이터셋 모두에서 강력한 성과를 달성하여 사전 계산된 광학 흐름 없이도 빠른 동작 인식을 검증한다.
  • VAP는 경량 매개변수로 여러 시간척도에 걸친 효과적인 장기 시간 모델링을 가능하게 하여 비디오 수준 표현을 향상시킨다.
  • 전통적인 광학 흐름 방법(TV-L1, FlowNet 계열)과 비교할 때 PA는 상당히 더 높은 속도와 경쟁력 있는 정확도를 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.