Skip to main content
QUICK REVIEW

[논문 리뷰] Deep Temporal Linear Encoding Networks

Ali Diba, Vivek Sharma|arXiv (Cornell University)|2016. 11. 21.
Human Pose and Action Recognition참고 문헌 39인용 수 12
한 줄 요약

이 논문은 전체 영상의 시공간적 특징을 희박한 영상 세그먼트로 인코딩하여 압축되고 강건한 표현으로 통합하는 새로운 엔드투엔드 딥러닝 레이어인 시간선형인코딩(TLE)을 제안한다. TLE는 2D 및 3D 컨볼루션 네트워크를 개선하여 행동 인식 성능을 향상시키며, HMDB51(71.1%)과 UCF101(95.6%)에서 최고 성능을 기록함과 동시에 효율성과 특징 학습 능력이 향상된다.

ABSTRACT

The CNN-encoding of features from entire videos for the representation of human actions has rarely been addressed. Instead, CNN work has focused on approaches to fuse spatial and temporal networks, but these were typically limited to processing shorter sequences. We present a new video representation, called temporal linear encoding (TLE) and embedded inside of CNNs as a new layer, which captures the appearance and motion throughout entire videos. It encodes this aggregated information into a robust video feature representation, via end-to-end learning. Advantages of TLEs are: (a) they encode the entire video into a compact feature representation, learning the semantics and a discriminative feature space; (b) they are applicable to all kinds of networks like 2D and 3D CNNs for video classification; and (c) they model feature interactions in a more expressive way and without loss of information. We conduct experiments on two challenging human action datasets: HMDB51 and UCF101. The experiments show that TLE outperforms current state-of-the-art methods on both datasets.

연구 동기 및 목표

  • 전체 영상에 걸쳐 장거리 시간적 의존성을 모델링하는 데에 한계를 가진 기존의 컨볼루션 네트워크를 해결한다.
  • 밀도 높은 시간 샘플링이나 고정 길이 클립에 의존하는 전통적인 이중 스트림 및 3D 컨볼루션 네트워크에서 발생하는 계산 비효율성과 정보 손실 문제를 해결한다.
  • 모델 파라미터 수를 늘리지 않고도 특징 표현을 향상시키는 일반적인 목적의 경량 인코딩 레이어를 개발한다.
  • 장기간의 시퀀스에 걸쳐 외관과 운동 역학을 모두 포괄하는 글로벌 영상 수준의 특징을 엔드투엔드 학습으로 가능하게 한다.
  • 다양한 아키텍처와 데이터셋, 특히 HMDB51 및 UCF101에서의 효과성과 유연성을 입증한다.

제안 방법

  • TLE는 전체 영상 기간 동안 분포된 고정된 수의 희박한 영상 세그먼트(프레임 또는 클립)를 처리한다.
  • 각 세그먼트는 공유 가중치를 가진 컨볼루션 네트워크(2D 또는 3D)를 통해 처리되어 공간적 및 시간적 특징을 추출한다.
  • 모든 세그먼트의 특징은 학습 가능한 선형 투영 레이어를 통해 집계되어 압축된 글로벌 영상 임베딩을 형성한다.
  • 이중 또는 최대 풀링 기반 융합을 사용하여 낮은 차원의 공간에서 세그먼트 간 특징 상호작용을 모델링한다.
  • TLE는 기존의 컨볼루션 네트워크 아키텍처 내에서 미분 가능한 레이어로 통합되어 역전파를 통한 엔드투엔드 학습이 가능하다.
  • 사전 훈련된 공간 스트림을 추가함으로써 장소365와 같은 추가 데이터 스트림과의 융합을 지원한다.
Figure 1 : Temporal linear encoding for video classification. Given several segments of an entire video, be it either a number of frames or a number of clips, the model builds a compact video representation from the spatial and temporal cues they contain, through end-to-end learning. The ConvNets ap
Figure 1 : Temporal linear encoding for video classification. Given several segments of an entire video, be it either a number of frames or a number of clips, the model builds a compact video representation from the spatial and temporal cues they contain, through end-to-end learning. The ConvNets ap

실험 결과

연구 질문

  • RQ1학습 가능한 압축된 시간 인코딩 레이어는 엔드투엔드 컨볼루션 네트워크에서 글로벌 영상 표현 학습을 향상시킬 수 있는가?
  • RQ2TLE는 표준 행동 인식 벤치마크에서 수작업 특징(iDT+FV)과 최신 딥러닝 방법에 비해 어떻게 비교되는가?
  • RQ3TLE는 두 스트림 및 C3D 네트워크를 포함한 다양한 네트워크 아키텍처에서 성능 향상에 얼마나 기여하는가?
  • RQ4장소 정보와 같은 보조 정보를 효과적으로 통합하여 행동 인식 정확도를 더 높일 수 있는가?
  • RQ5TLE는 풀링이나 잘라내기로 인한 정보 손실을 피하면서도 모델 복잡도를 줄이며도 여전히 분류 능력을 유지하는가?

주요 결과

  • TLE는 UCF101에서 최고 성능인 95.6%와 HMDB51에서 71.1%를 기록하여 iDT+FV 및 이중 스트림 네트워크와 같은 이전 방법들을 능가한다.
  • TLE:Bilinear은 C3D 컨볼루션 네트워크의 성능을 UCF101에서 86.3%로 향상시키고, HMDB51에서는 60.3%로 높여 원본 C3D 및 이중 스트림 기반선보다 뛰어나다.
  • 장소365 사전 훈련된 특징를 통한 장소 컨텍스트 추가로 성능이 UCF101에서 83.8%로 향상되고, HMDB51에서는 63.6%로 상승하여 상호 보완적 특징 학습이 가능함을 입증한다.
  • 완전 연결층 대비 모델 파라미터를 크게 줄였지만 표현력 있는 특징 상호작용을 유지한다.
  • TLE는 계산적으로 효율적이며 강건하여 밀도 높은 프레임 처리 없이도 효과적인 장거리 시간 모델링이 가능하다.
  • TLE는 2D 및 3D 컨볼루션 네트워크를 포함한 다양한 아키텍처에 잘 일반화되며, 다른 순차적 데이터 스트림으로도 확장 가능하다.
Figure 2 : Our temporal linear encoding applied to the design of two-stream ConvNets [ 24 ] : spatial and temporal networks. The spatial network operates on RGB frames, and the temporal network operates on optical flow fields. The features maps from the spatial and temporal ConvNets for multiple suc
Figure 2 : Our temporal linear encoding applied to the design of two-stream ConvNets [ 24 ] : spatial and temporal networks. The spatial network operates on RGB frames, and the temporal network operates on optical flow fields. The features maps from the spatial and temporal ConvNets for multiple suc

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.