Skip to main content
QUICK REVIEW

[논문 리뷰] Pedestrian Action Anticipation using Contextual Feature Fusion in Stacked RNNs

Amir Rasouli, Iuliia Kotseruba|arXiv (Cornell University)|2020. 05. 13.
Autonomous Vehicle Technology and Safety인용 수 65
한 줄 요약

다섯 수준의 스택드 GRU 아키텍처(SF-GRU)를 도입하여 다중 모달 맥락 특징을 융합하고 보행자 횡단 예측에서 여러 RNN 기초 모델보다 우수한 정확도를 보인다.

ABSTRACT

One of the major challenges for autonomous vehicles in urban environments is to understand and predict other road users' actions, in particular, pedestrians at the point of crossing. The common approach to solving this problem is to use the motion history of the agents to predict their future trajectories. However, pedestrians exhibit highly variable actions most of which cannot be understood without visual observation of the pedestrians themselves and their surroundings. To this end, we propose a solution for the problem of pedestrian action anticipation at the point of crossing. Our approach uses a novel stacked RNN architecture in which information collected from various sources, both scene dynamics and visual features, is gradually fused into the network at different levels of processing. We show, via extensive empirical evaluations, that the proposed algorithm achieves a higher prediction accuracy compared to alternative recurrent network architectures. We conduct experiments to investigate the impact of the length of observation, time to event and types of features on the performance of the proposed method. Finally, we demonstrate how different data fusion strategies impact prediction accuracy.

연구 동기 및 목표

  • 트랙토리 기반 방법을 넘어서는 강건한 보행자 횡단 예측을 위해 전체 맥락 신호를 활용한다.
  • Appearance, context, pose, bounding box, 및 ego-vehicle speed를 점진적으로 통합하는 다층 피처 융합 아키텍처를 개발한다.
  • 관찰 길이, 시간-대-이벤트(TTE), 및 피처 융합 순서가 예측 성능에 미치는 영향을 평가한다.
  • 다른 RNN 아키텍처와 SF-GRU를 다양한 관찰 설정에서 비교한다.

제안 방법

  • 하위에서 상위로 다중 모달 피처를 점진적으로 융합하는 다층 GRU(SF-GRU) 5층을 제안한다.
  • 하위 계층을 보행자 외관 피처로 처리하고, 주변 맥락, 자세, 바운딩 박스, 및 자가 차량 속도를 점진적으로 융합한다.
  • x^t_0 = vc^t_p와 함께 y^t = {vc^t_s, vp^t, vb^t, vs^t}로 정의된 방식에서 각 레벨마다 입력을 연결(concatenate)하는 256개의 은닉 유닛 GRU를 사용한다.
  • Adam 옵티마이저로 이진 교차 엔트로피 손실을 사용해 60 에폭 동안 학습; 컨텍스트 및 포즈 피처를 미리 계산; 수평 반전 및 클래스 재샘플링으로 데이터 보강.
  • 시간-대-이벤트(TTE) 및 관찰 길이가 예측 성능에 미치는 영향을 평가한다; 피처 융합 순서가 정확도에 어떻게 영향을 미치는지 조사한다.

실험 결과

연구 질문

  • RQ1다중 모달 다층 융합 모델이 궤도 기반 방법을 넘어 보행자 횡단 예측을 개선할 수 있는가?
  • RQ2관찰 길이와 시간-대-이벤트가 횡단 예측 정확도에 미치는 영향은 무엇인가?
  • RQ3다양한 데이터 융합 순서와 피처 입력이 SF-GRU 성능에 어떤 영향을 미치는가?
  • RQ4Appearance, context, pose, bounding box, 및 ego-vehicle speed 피처의 어떤 조합이 최적의 정확도를 제공하는가?

주요 결과

  • SF-GRU가 평가된 모델 중에서 0.844의 정확도와 0.829의 AUC를 달성했다.
  • SF-GRU가 대부분의 지표에서 Static, GRU, M-GRU, H-GRU 기초 모델보다 우수한 F1 0.721, 정밀도 0.657, 재현율 0.800을 달성했다.
  • 모든 정보 소스(C_p, C_s, P, B, S)를 사용하는 것이 부분 피처 세트보다 우수한 성능을 낳는다.
  • 외관과 주변 맥 context를 해체하고 섬세한 피처 융합 순서가 정확도(일부 구성에서 정확도 최대 9%, 재현율 10%, 정밀도 15% 향상)에 큰 영향을 미친다.
  • 관찰 시간이 길수록 근단(0–1초) 및 원거리(3초) 영역에서 주로 도움이 되며, 장면의 동적 노이즈로 인해 중간 구간에서 혼합 효과를 보인다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.