Skip to main content
QUICK REVIEW

[論文レビュー] Intelligent Traffic Signal Control: Using Reinforcement Learning with Partial Detection

Rusheng Zhang, Akihiro Ishikawa|arXiv (Cornell University)|Jul 4, 2018
Traffic control and management被引用数 14
ひとこと要約

本稿では、DSRC対応のV2I通信を備えた車両のみが観測される部分的に検出可能なインテリジェント交通システム(ITS)を想定し、強化学習(RL)に基づく交通信号制御システムを提案する。部分的観測と低検出率を扱うためにRLを活用することで、20%の検出率でさえも平均車両待機時間を顕著に短縮し、多様な交通状況およびネットワーク環境においても頑健な性能を示す。

ABSTRACT

Intelligent Transportation Systems (ITS) have attracted the attention of researchers and the general public alike as a means to alleviate traffic congestion. Recently, the maturity of wireless technology has enabled a cost-efficient way to achieve ITS by detecting vehicles using Vehicle to Infrastructure (V2I) communications. Traditional ITS algorithms, in most cases, assume that every vehicle is observed, such as by a camera or a loop detector, but a V2I implementation would detect only those vehicles with wireless communications capability. We examine a family of transportation systems, which we will refer to as `Partially Detected Intelligent Transportation Systems'. An algorithm that can act well under a small detection rate is highly desirable due to gradual penetration rates of the underlying wireless technologies such as Dedicated Short Range Communications (DSRC) technology. Artificial Intelligence (AI) techniques for Reinforcement Learning (RL) are suitable tools for finding such an algorithm due to utilizing varied inputs and not requiring explicit analytic understanding or modeling of the underlying system dynamics. In this paper, we report a RL algorithm for partially observable ITS based on DSRC. The performance of this system is studied under different car flows, detection rates, and topologies of the road network. Our system is able to efficiently reduce the average waiting time of vehicles at an intersection, even with a low detection rate.

研究の動機と目的

  • 無線技術の浸透が限定的であるため、一部の車両のみが観測される部分的検出可能なITSにおける交通信号制御の課題に対処すること。
  • 初期段階のDSRC導入に見られる低検出率環境でも良好に機能する知能的な制御アルゴリズムを開発すること。
  • 交通ダイナミクスの完全な知識や環境の明示的モデリングを必要としないシステムを設計すること。
  • varying traffic flows, detection rates, and road network topologies におけるRLベースのコントローラーの性能を評価すること。

提案手法

  • 交通ダイナミクスの明示的モデリングを必要とせず、最適な信号タイミングポリシーを学習するため、強化学習(RL)を採用する。
  • DSRC搭載車両からの状態情報のみを用いるため、部分的観測下でも動作可能なRLエージェントを設計する。
  • 観測された車両状態を最適な信号フェーズ行動にマッピングする価値ベースのRLアプローチ(例:Q学習やDQNに類似)を用いる。
  • 異なる検出率およびネットワークトポロジーを有するシミュレーテッド交通環境を用いてエージェントを訓練する。
  • 部分的な車両検出を統合して信号制御に役立てる行動可能な交通推定値を生成する状態表現を組み込む。
  • 交差点における平均車両遅延を最小化するように、RL報酬関数を最適化する。

実験結果

リサーチクエスチョン

  • RQ1DSRC搭載車両の検出率が低下するにつれて、RLベースの交通信号コントローラーの性能はどの程度低下するか?
  • RQ2一部の車両のみが観測可能である状況でも、RLエージェントは効果的な信号制御ポリシーを学習できるか?
  • RQ3部分的検出下で、異なる交通フロー水準および道路ネットワーク構成において、システムの性能はいかがなっているか?
  • RQ4RLコントローラーが平均車両待機時間を有意義に短縮できる最小の検出率はどの程度か?

主な発見

  • RLベースのコントローラーは、低検出率でも交差点における平均車両待機時間を顕著に短縮する。
  • 20%という低検出率でも、依然として優れた性能を維持し、部分的観測に対する耐性を示す。
  • 完全観測の仮定が成り立たない部分的検出環境においても、従来手法を上回る性能を発揮する。
  • 多様な交通状況および道路ネットワークトポロジーにわたり、性能が頑健である。
  • 明示的なシステムモデリングを必要とせず、限られた入力情報のみで最適な信号タイミング意思決定を効果的に学習する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。