[論文レビュー] PSI: A Pedestrian Behavior Dataset for Socially Intelligent Autonomous Car
本稿では、時間的動的歩行者横断意図のアノテーションと人間が生成した推論説明を備えた、社会的知能を持つ自動運転車両のための新規ベンチマークであるPSIデータセットを紹介する。本稿では、これらの認知的ラベルを活用する説明可能な歩行者軌道予測モデル(eP2P)を提案し、意図予測(F1: 0.66)および説明性と不一致検出性を向上させた軌道予測で最先端の性能を達成した。
Prediction of pedestrian behavior is critical for fully autonomous vehicles to drive in busy city streets safely and efficiently. The future autonomous cars need to fit into mixed conditions with not only technical but also social capabilities. As more algorithms and datasets have been developed to predict pedestrian behaviors, these efforts lack the benchmark labels and the capability to estimate the temporal-dynamic intent changes of the pedestrians, provide explanations of the interaction scenes, and support algorithms with social intelligence. This paper proposes and shares another benchmark dataset called the IUPUI-CSRC Pedestrian Situated Intent (PSI) data with two innovative labels besides comprehensive computer vision labels. The first novel label is the dynamic intent changes for the pedestrians to cross in front of the ego-vehicle, achieved from 24 drivers with diverse backgrounds. The second one is the text-based explanations of the driver reasoning process when estimating pedestrian intents and predicting their behaviors during the interaction period. These innovative labels can enable several computer vision tasks, including pedestrian intent/behavior prediction, vehicle-pedestrian interaction segmentation, and video-to-language mapping for explainable algorithms. The released dataset can fundamentally improve the development of pedestrian behavior prediction models and develop socially intelligent autonomous cars to interact with pedestrians efficiently. The dataset has been evaluated with different tasks and is released to the public to access.
研究の動機と目的
- 自動運転のための既存データセットにおける動的で状況に依存する歩行者意図アノテーションの不足に対処する。
- 車両・歩行者相互作用の過程でドライバーが生成したテキストベースの推論説明を導入し、アルゴリズムの説明可能性を向上させる。
- 歩行者意図、軌道、明示的な推論説明を同時に予測する統合モデルの開発。
- 意図の動的変化と社会的合意形成の兆候をモデル化することで、より社会的知能を持つ自動運転車両の行動を実現する。
- 説明可能な歩行者行動予測分野の研究を促進するため、公開可能なベンチマークデータセットを提供する。
提案手法
- 24名の多様な人間ドライバーが時間的動的歩行者横断意図をアノテートした、110本の実世界の都市走行映像クリップを収集・アノテートした。
- 各意図推定に対して、相互作用のタイムラインと同期したドライバーの推論説明のテキストを収集した。
- 歩行者意図とその対応する説明を条件として用いるエンドツーエンドの説明可能な歩行者軌道予測モデル(eP2P)を設計した。
- 動的意図ラベルと自然言語の説明を含む、動画特徴量、ボクシングボックス、マルチモーダル入力を用いてeP2Pモデルを学習した。
- 意図分類、軌道予測、説明生成の3つを同時に最適化するマルチタスク学習フレームワークを実装した。
- 不確実な予測を特定する不一致検出を統合し、複雑な相互作用における安全性と解釈可能性を向上させた。
実験結果
リサーチクエスチョン
- RQ1車両・歩行者相互作用の過程で、歩行者意図を状況依存の動的変数としてどのように定義・アノテートできるか。
- RQ2人間が生成した推論説明は、歩行者行動予測モデルの性能と解釈可能性をどの程度向上させ得るか。
- RQ3統合モデルが1つのフレームワーク内で歩行者意図、軌道、自然言語の説明を効果的に予測できるか。
- RQ4動的意図と説明アノテーションを組み込むことで、ベースラインと比較して軌道予測のロバスト性と一般化性能がどのように向上するか。
- RQ5説明可能なモデリングは、自動運転車両の意思決定における信頼性とユーザーの信頼にどのような影響を与えるか。
主な発見
- 提案されたeP2Pモデルは、二値意図分類タスクにおいて、バランス精度0.67、F1スコア0.66を達成し、PSIデータセット上でPIEベースラインモデルを上回った。
- eP2Pモデルは、1.5秒後のLSTMベースラインと比較して、平均移動誤差(ADE)を15.8%、最終移動誤差(FDE)を12.4%削減した。
- モデルは関連性のある自然言語の説明を効果的に生成できており、生成された説明の68%が正解の説明と内容的に一致した。
- PSIデータセットは、24名のアノテータ間で意図アノテーションに顕著なばらつきを示しており、現実世界の歩行者相互作用の複雑さと動的性質を浮き彫りにした。
- PIEベースラインモデルは、PSIデータセットに適用された際、性能が著しく低下した(F1: 0.79 → 0.66)ことから、動的で文脈に依存する意図シナリオへの一般化能力に限界があることが示された。
- 説明生成と不一致検出の統合により、モデルは不確実な予測を特定でき、高リスクな状況における安全性が向上した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。