[논문 리뷰] Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation
이 논문은 장기적인 2D 자세 시계열에서 전역적이고 국소적인 시간적 맥락을 계층적으로 통합하기 위해 피드포워드 네트워크에 스트라이드 컨볼루션을 활용하는 새로운 Transformer 기반 아키텍처인 Strided Transformer를 제안한다. 이는 계산량을 크게 줄이며 3D 자세 추정 정확도를 향상시킨다. 이 방법은 전체-단일 감독 스킴을 통해 더 적은 파라미터로 Human3.6M 및 HumanEva-I에서 최신 기술 수준(SOTA) 성능을 달성하며, 시간적 매끄러움이 향상된다.
Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we propose an improved Transformer-based architecture, called Strided Transformer, which simply and effectively lifts a long sequence of 2D joint locations to a single 3D pose. Specifically, a Vanilla Transformer Encoder (VTE) is adopted to model long-range dependencies of 2D pose sequences. To reduce the redundancy of the sequence, fully-connected layers in the feed-forward network of VTE are replaced with strided convolutions to progressively shrink the sequence length and aggregate information from local contexts. The modified VTE is termed as Strided Transformer Encoder (STE), which is built upon the outputs of VTE. STE not only effectively aggregates long-range information to a single-vector representation in a hierarchical global and local fashion, but also significantly reduces the computation cost. Furthermore, a full-to-single supervision scheme is designed at both full sequence and single target frame scales applied to the outputs of VTE and STE, respectively. This scheme imposes extra temporal smoothness constraints in conjunction with the single target frame supervision and hence helps produce smoother and more accurate 3D poses. The proposed Strided Transformer is evaluated on two challenging benchmark datasets, Human3.6M and HumanEva-I, and achieves state-of-the-art results with fewer parameters. Code and models are available at \url{https://github.com/Vegetebird/StridedTransformer-Pose3D}.
연구 동기 및 목표
- 2D 자세 시계열에서 3D 인간 자세 추정을 위해 장기적인 시간적 의존성을 효율적으로 모델링하는 데 도전한다.
- 장기적인 영상 시계열을 처리할 때 표준 Transformer의 이차적 계산 비용을 줄인다.
- 풀링으로 인한 정보 손실을 방지하기 위해 시퀀스 길이를 점진적으로 축소하면서도 세밀한 국소적 특징을 유지한다.
- 전체-단일 감독 스킴을 도입하여 추정 정확도와 시간적 매끄러움을 향상시킨다.
제안 방법
- 기본 Transformer 인코더(VTE)의 피드포워드 네트워크에서 완전 연결층을 스트라이드 컨볼루션으로 대체하여 시퀀스 길이를 점진적으로 감소시키고 국소적 및 전역적 맥락을 통합한다.
- VTE 출력을 처리하여 2D 자세 시계열의 압축되고 계층적인 표현을 생성하는 Strided Transformer 인코더(Ste)를 도입한다.
- 전체 시퀀스 및 단일 타겟 프레임 수준에서 감독을 적용하는 전체-단일 감독 스킴을 설계하여 시간적 일관성을 향상시킨다.
- VTE에서 다중 헤드 자기주의 어텐션을 사용해 모든 프레임 간의 장기적 의존성을 모델링하고, STE의 스트라이드 컨볼루션은 국소 패턴 통합에 집중한다.
- 정확도와 강건성을 향상시키기 위해 시퀀스 수준 및 프레임 수준의 손실을 조합하여 모델을 엔드 투 엔드로 훈련시킨다.
- VTE와 STE가 함께 최적화되어 상보적인 표현을 학습할 수 있도록 이중 단계 훈련 전략을 활용한다.
실험 결과
연구 질문
- RQ1Transformer의 피드포워드 네트워크에 스트라이드 컨볼루션을 적용하여 시퀀스 길이를 효과적으로 줄일 수 있을까? 이는 3D 자세 추정을 위한 핵심 시간적 정보를 유지하는가?
- RQ2Strided Transformer의 계층적 전역 및 국소 특징 통합 방식이 표준 Transformer보다 3D 자세 추정 정확도를 어떻게 향상시키는가?
- RQ3전체-단일 감독 스킴이 예측된 3D 자세의 시간적 매끄러움과 표현 품질을 어느 정도 향상시키는가?
- RQ4제안된 방법이 기존의 Transformer 기반 접근 방식보다 더 적은 파라미터와 낮은 계산 비용으로 최신 기술 수준 성능을 달성할 수 있는가?
주요 결과
- 제안된 Strided Transformer는 Human3.6M 및 HumanEva-I에서 최신 기술 수준 성능을 달성하였으며, Human3.6M에서 평균 MPJPE는 40.2 mm, HumanEva-I에서 47.6 mm를 기록하였다.
- FFN를 스트라이드 컨볼루션으로 대체함으로써 계산 비용을 크게 줄였고, 표준 Transformer와 비교해 정확도를 유지하거나 향상시켰다.
- 전체-단일 감독 스킴은 풀링 기반 베이스라인 대비 MPJPE를 0.4 mm 감소시켜 시간적 매끄러움 향상 효과를 입증하였다.
- VTE 모듈을 제거하면 MPJPE가 1.1 mm 증가하여, 장기적 의존성 모델링에서 VTE의 중요성을 입증하였다.
- STE 모듈을 제거하면 MPJPE가 47.6 mm로 증가하여, 계층적 국소 및 전역 특징 통합에서 STE의 핵심적 역할을 확인하였다.
- 정성적 결과에서는 빠른 동작과 부분적 가림이 있는 실제 환경 영상에서도 모델이 현실적이고 구조적으로 타당한 3D 자세를 생성하는 것으로 나타났다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.