Skip to main content
QUICK REVIEW

[論文レビュー] When Do Drivers Concentrate? Attention-based Driver Behavior Modeling With Deep Reinforcement Learning

Xingbo Fu, Gao, Feng|arXiv (Cornell University)|Feb 26, 2020
Autonomous Vehicle Technology and Safety参考文献 14被引用数 4
ひとこと要約

本稿では、実際の軌道データを用いて、車両追従シナリオにおけるドライバーの時間的注意割り当てをモデル化する、注意を強化したアクタ・クリティック強化学習フレームワークであるATD3を提案する。アクタネットワークに時間的注意メカニズムを統合し、価値推定にTD3を採用することで、7つのベースラインを上回るドライバー行動の予測性能を達成し、急激な速度低下時には注意が均一な状態から最近の観測にシフトすることが明らかになった。これにより、注意の欠落パターンに関する新たな知見が得られた。

ABSTRACT

Driver distraction a significant risk to driving safety. Apart from spatial domain, research on temporal inattention is also necessary. This paper aims to figure out the pattern of drivers' temporal attention allocation. In this paper, we propose an actor-critic method - Attention-based Twin Delayed Deep Deterministic policy gradient (ATD3) algorithm to approximate a driver' s action according to observations and measure the driver' s attention allocation for consecutive time steps in car-following model. Considering reaction time, we construct the attention mechanism in the actor network to capture temporal dependencies of consecutive observations. In the critic network, we employ Twin Delayed Deep Deterministic policy gradient algorithm (TD3) to address overestimated value estimates persisting in the actor-critic algorithm. We conduct experiments on real-world vehicle trajectory datasets and show that the accuracy of our proposed approach outperforms seven baseline algorithms. Moreover, the results reveal that the attention of the drivers in smooth vehicles is uniformly distributed in previous observations while they keep their attention to recent observations when sudden decreases of relative speeds occur. This study is the first contribution to drivers' temporal attention and provides scientific support for safety measures in transportation systems from the perspective of data mining.

研究の動機と目的

  • 動的ドライブシナリオ、特に相対速度の変化が生じる際のドライバーの時間的注意割り当ての仕組みをモデル化すること。
  • ドライブセーフティ分野における空間的注意の欠落にとどまらない、時間的注意の欠落に関する研究ギャップを埋めること。
  • 注意メカニズムを用いてドライバー行動の時間的依存性を捉える、深層強化学習フレームワークの開発。
  • 注意ダイナミクスをモデル化することで、車両追従タスクにおける行動予測の正確性を向上させること。

提案手法

  • 注意メカニズムとTwin Delayed Deep Deterministic Policy Gradient (TD3)を組み合わせた、アクタ・クリティックアルゴリズムであるATD3を提案する。
  • アクタネットワークに時間的注意メカニズムを統合し、過去の観測を現在の意思決定に対する関連性に基づいて重み付けする。
  • 価値関数学習における過剰推定バイアスを低減するために、クリティックネットワークでTD3を採用する。
  • 実世界の車両軌道データセットを用いて、車両追従タスクにおけるドライバー行動の予測を目的として、モデルをエンドツーエンドで学習する。
  • 連続する観測に注目することで反応時間のモデル化を実現し、ドライバー行動の時間的依存性を捉える。
  • アクタが注意を向けた観測に基づいて行動を生成し、クリティックが行動価値を評価する二重ネットワークアーキテクチャを採用する。

実験結果

リサーチクエスチョン

  • RQ1ドライバーは車両追従タスクにおいて、時間的にどのように注意を配分しているか?
  • RQ2相対速度が急激に変化する際、ドライバーの注意割り当てにどのようなパターンが現れるか?
  • RQ3注意メカニズムを備えた深層強化学習は、ベースラインモデルと比較してドライバー行動予測の正確性を向上させることができるか?
  • RQ4時間的注意の組み込みが、ドライバー行動ダイナミクスを捉えるモデルの能力にどのように影響するか?

主な発見

  • 提案されたATD3モデルは、実世界の軌道データセットにおいて7つのベースラインアルゴリズムを上回る高い予測精度を達成した。
  • 滑らかな車両を追従している際、ドライバーは過去の観測に対して均等に注意を割り当てることが明らかになった。
  • 相対速度が急激に低下する際、ドライバーは注意をより最近の観測にシフトすることが観察された。
  • 注意メカニズムはドライバー行動の時間的ダイナミクスを効果的に捉えており、文脈依存的な注意パターンを明らかにした。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。