[논문 리뷰] Highly Efficient 3D Human Pose Tracking from Events with Spiking Spatiotemporal Transformer
논문은 이벤트 스트림만으로 3D 인체 포즈를 추적하기 위해 스파이킹 신경망(SNN)을 활용한 엔드투엔드 희소 딥러닝 접근법을 제시하며, Spiking Spatiotemporal Transformer와 대규모 합성 SynEventHPD 데이터셋을 특징으로 하고, 최첨단 방법들을 상회하는 동시에 FLOPs를 크게 감소시킨다.
Event camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatiotemporal Transformer, which enables bi-directional spatiotemporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a large-scale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN.
연구 동기 및 목표
- 3D 인체 포즈 추적을 gray-scale 프레임 없이 이벤트 카메라 데이터만으로 수행한다.
- Spiking Spatiotemporal Transformer를 활용한 엔드-투-엔드 SNN 아키텍처를 개발하여 양방향 시간 융합을 달성한다.
- 다양한 모션을 지원하기 위해 대규모 합성 이벤트 기반 데이터셋 SynEventHPD를 도입한다.
- 최신 상태의 ANN 및 SNN 기준선보다 더 우수한 성능과 상당한 계산 효율성을 입증한다.
제안 방법
- 이벤트 스트림을 시간 정보를 보존하도록 이벤트 보셀 그리드의 시퀀스로 전처리한다.
- voxel grids에서 포즈 스파이크 특징을 추출하기 위해 SNN 백본으로 SEW-ResNet을 사용한다.
- 스파이킹 스파이타임포럴 트랜스포머를 도입하여 스파이크 특징의 시간적 융합을 위한 양방향 어텐션을 구현한다.
- 2D 풀링된 스파이크 특징으로부터 세 개의 병렬 선형 계층을 통해 SMPL 포즈 및 모양 파라미터를 회귀한다.
- 포즈, 모양, 3D 및 2D 관절 손실로 엔드-투-엔드 학습하여 시간 정합된 3D 메쉬를 출력한다.

실험 결과
연구 질문
- RQ13D 인체 포즈 추적을 events 만으로, gray-scale 프레임 없이 엔드-투-엔드로 수행할 수 있는가?
- RQ2SNN에서 포즈 추적을 위한 스파이킹 시공간 어텐션 메커니즘이 양방향 시간 융합을 가능하게 하는가?
- RQ3완전한 SNN 기반 접근법은 이벤트 기반 포즈 추적에서 정확도와 계산 측면에서 ANN/SNN 하이브리드와 어떻게 비교되는가?
- RQ4대규모 합성 이벤트 기반 데이터셋이 이벤트 주도 포즈 추적의 일반화에 어떤 영향을 미치는가?
주요 결과
- 해당 방법은 이벤트 기반 3D 포즈 추적에서 최신 ANN 및 기준선 SNN 방법들보다 우수한 성능을 달성한다.
- 제안하는 방법은 SOTA 방법에 비해 FLOPs를 약 80% 감소시킨다.
- Spiking Spatiotemporal Transformer는 시간 초반의 포즈 추정을 개선하기 위한 양방향 정보 흐름을 가능하게 한다.
- 다양한 모션 데이터셋에서 총 45.72 시간의 이벤트 스트림으로 구성된 대규모 합성 데이터셋 SynEventHPD를 도입한다.
- 파이프라인은 시간에 따라 SMPL 파라미터(beta, theta)와 전역 변환을 출력하여 각 시간 스텝에 대해 3D 메쉬를 형성한다.
![Figure 2: Pipeline of our sparse deep learning approach . It contains four main sections: (i) Preprcessing in Sec. 4.1 converts a stream of events into a sequence of event voxel grids of the same temporal length. (ii) SEW-ResNet [ 24 ] , introduced in Sec. 4.2 , is used as backbone to extract pose s](https://ar5iv.labs.arxiv.org/html/2303.09681/assets/x2.png)
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.