[논문 리뷰] MAR: Masked Autoencoders for Efficient Action Recognition
이 논문은 비디오 동작 인식에서 비전 트랜스포머(ViTs)를 위한 계산적으로 효율적인 학습 방법인 마스킹 동작 인식(MAR)을 제안한다. 공간시계열 상관관계를 유지하기 위해 셀 러닝 마스킹을 적용하고, 재구성과 분류 특징 간의 의미적 갭을 메우기 위해 브리지 분류기를 도입함으로써, MAR은 ViT의 계산량을 53% 감소시키면서도 표준 학습 방식을 능가한다. 예를 들어, MAR을 사용해 훈련한 ViT-Large는 표준 학습 방식으로 훈련한 ViT-Huge보다 정확도는 높지만 계산 비용은 오직 14.5%에 불과하다.
Standard approaches for video recognition usually operate on the full input videos, which is inefficient due to the widely present spatio-temporal redundancy in videos. Recent progress in masked video modelling, i.e., VideoMAE, has shown the ability of vanilla Vision Transformers (ViT) to complement spatio-temporal contexts given only limited visual contents. Inspired by this, we propose propose Masked Action Recognition (MAR), which reduces the redundant computation by discarding a proportion of patches and operating only on a part of the videos. MAR contains the following two indispensable components: cell running masking and bridging classifier. Specifically, to enable the ViT to perceive the details beyond the visible patches easily, cell running masking is presented to preserve the spatio-temporal correlations in videos, which ensures the patches at the same spatial location can be observed in turn for easy reconstructions. Additionally, we notice that, although the partially observed features can reconstruct semantically explicit invisible patches, they fail to achieve accurate classification. To address this, a bridging classifier is proposed to bridge the semantic gap between the ViT encoded features for reconstruction and the features specialized for classification. Our proposed MAR reduces the computational cost of ViT by 53% and extensive experiments show that MAR consistently outperforms existing ViT models with a notable margin. Especially, we found a ViT-Large trained by MAR outperforms the ViT-Huge trained by a standard training scheme by convincing margins on both Kinetics-400 and Something-Something v2 datasets, while our computation overhead of ViT-Large is only 14.5% of ViT-Huge.
연구 동기 및 목표
- 전체 비디오 프레임을 처리하는 표준 비디오 동작 인식 방식이 광범위한 공간시계열 중복성에도 불구하고 효율성이 떨어지는 문제를 해결하기 위해.
- 성능을 저하시키지 않은 채 비전 트랜스포머(ViT) 기반의 동작 인식에서 계산 비용을 줄이기 위해.
- 마스킹된 패치의 부분집합만을 사용하여 ViT가 손실된 시각적 콘텐츠를 효과적으로 재구성하고 동작를 분류할 수 있도록 하기 위해.
- 재구성에 사용되는 특징과 분류 전용 특징 간의 의미적 갭을 메우기 위해.
제안 방법
- 셀 러닝 마스킹을 도입하여 공간시계열으로 겹치는 마스크를 생성함으로써, 동일한 공간 위치에 있는 패치들이 프레임 간에 순차적으로 관측되도록 하여 시간적 일관성을 유지한다.
- 마스킹 전략은 비디오 패치의 50~75%를 체계적으로 제거하면서도 재구성에 강력한 공간시계열적 맥락을 유지한다.
- 가벼운 브리지 분류기를 추가하여 ViT에 의해 인코딩된 재구성 특징의 의미 표현을 분류에 최적화된 특징과 정렬한다.
- 사전 훈련된 VideoMAE 인코더를 사용하여 종료형(end-to-end)으로 작동하며, 동작 인식을 위해 브리지 분류기와 헤드만 미세조정한다.
- 이 방법은 Kinetics-400 및 Something-Something v2 데이터셋에서 ViT-B, ViT-Large, ViT-Huge 모델에 적용된다.
- 마스킹된 입력을 사용해 훈련하고, 표준 동작 인식 벤치마크에서 평가한다.
실험 결과
연구 질문
- RQ1마스킹된 학습 방식은 ViT 기반 비디오 동작 인식에서 계산 비용을 줄일 수 있을까? 이때 정확도는 유지되거나 향상될 수 있는가?
- RQ2비디오의 공간시계열 중복성을 어떻게 활용하여, 가시 토큰의 부분집합만을 사용해 손실된 패치를 효과적으로 재구성할 수 있는가?
- RQ3전용 브리지 분류기가 재구성에 민감한 특징과 분류 전용 표현 간의 정렬에 미치는 영향은 무엇인가?
- RQ4MAR은 표준 학습 방식에 비해 훨씬 낮은 계산 비용으로 최신 기술 수준의 성능을 달성할 수 있는가?
- RQ5이 방법은 다양한 모델 크기와 규모가 다른 다양한 비디오 데이터셋에 일반화되는가?
주요 결과
- MAR는 표준 학습 방식 대비 ViT의 계산 비용을 53% 감소시켰고, Kinetics-400과 Something-Something v2 양쪽 모두에서 더 높은 성능를 달성했다.
- MAR로 훈련한 ViT-Large는 Kinetics-400에서 표준 학습 방식으로 훈련한 ViT-Huge보다 0.2% 높은 정확도를 기록했으며, Something-Something v2에서도 동일한 성과를 보였고, 계산 비용은 오직 GFLOPs의 14.5%에 불과했다.
- Kinetics-400에서 MAR을 사용한 ViT-Large는 131 GFLOPs에서 83.9%의 정확도를 달성했고, 표준 학습 방식의 ViT-Base(81.3%)보다 계산 비용을 27% 줄였다.
- Something-Something v2에서 벽시계 훈련 시간이 50% 감소하여, 표준 방식의 11.8시간에서 MAR의 5.9시간으로 줄었다.
- 유사한 계산 비용 예산 내에서 MAR은 최신 기술 수준의 성능를 달성했으며, 더 큰 데이터셋으로 사전 훈련된 방법들조차도 뛰어넘었다.
- UCF101 및 HMDB51와 같은 작은 데이터셋에서도 MAR은 경쟁력 있거나 우수한 성능를 유지하여 강력한 일반화 능력을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.