Skip to main content
QUICK REVIEW

[论文解读] Multi-Agent Reinforcement Learning for Dynamic Routing Games: A Unified Paradigm.

Zhenyu Shou, Xuan Di|arXiv (Cornell University)|Nov 22, 2020
Transportation Planning and Optimization参考文献 63被引用 8
一句话总结

本文提出了一种多智能体强化学习(MARL)范式,用于建模交通网络中自私智能体的动态路径选择行为,通过奖励工程统一了动态用户均衡(DUE)与动态系统最优(DSO)结果。结果表明,最优收费(Braess网络中≥25)与信号控制(Boulevard上4秒时移)可最小化平均行程时间并避免Braess悖论。

ABSTRACT

This paper aims to develop a unified paradigm that models one's learning behavior and the system's equilibrating processes in a routing game among atomic selfish agents. Such a paradigm can assist policymakers in devising optimal operational and planning countermeasures under both normal and abnormal circumstances. To this end, a multi-agent reinforcement learning (MARL) paradigm is proposed in which each agent learns and updates her own en-route path choice policy while interacting with others on transportation networks. This paradigm is shown to generalize the classical notion of dynamic user equilibrium (DUE) to model-free and data-driven scenarios. We also illustrate that the equilibrium outcomes computed from our developed MARL paradigm coincide with DUE and dynamic system optimal (DSO), respectively, when rewards are set differently. In addition, with the goal to optimize some systematic objective (e.g., overall traffic condition) of city planners, we formulate a bilevel optimization problem with the upper level as city planners and the lower level as a multi-agent system where each rational and selfish traveler aims to minimize her travel cost. We demonstrate the effect of two administrative measures, namely tolling and signal control, on the behavior of travelers and show that the systematic objective of city planners can be optimized by a proper control. The results show that on the Braess network, the optimal toll charge on the central link is greater or equal to 25, with which the average travel time of selfish agents is minimized and the emergence of Braess paradox could be avoided. In a large-sized real-world road network with 69 nodes and 166 links, the optimal offset for signal control on Broadway is derived as 4 seconds, with which the average travel time of all controllable agents is minimized.

研究动机与目标

  • 开发一种统一的MARL范式,以建模动态路径选择博弈中的个体学习行为与系统级均衡化。
  • 通过MARL将经典DUE推广至无模型、数据驱动的场景。
  • 通过收费与信号配时等行政调控手段,优化城市规划者系统性目标(如最小化总体行程时间)。
  • 研究政策干预如何影响现实与合成网络中自私智能体的行为及整体系统效率。
  • 证明在特定奖励设置下,MARL的均衡结果与DUE和DSO一致。

提出的方法

  • 构建双层优化问题:上层为城市规划者,下层为使用MARL学习路径策略的自利智能体。
  • 采用MARL,每个智能体通过在网络中的交互独立学习并更新其路径选择策略。
  • 使用奖励塑造技术,引导MARL结果趋向DUE(个体最小化)或DSO(系统最优)。
  • 将收费与信号控制作为行政调控手段,以影响智能体行为与系统性能。
  • 在Braess网络与一个69节点、166条边的真实世界网络上验证该框架。
  • 通过仿真评估,推导出能最小化平均行程时间的最优控制参数(收费与信号时移)。

实验结果

研究问题

  • RQ1MARL框架能否统一建模动态路径选择博弈中的个体学习与系统均衡?
  • RQ2MARL中不同的奖励函数如何影响收敛至DUE或DSO结果?
  • RQ3在Braess网络中,中央路段的最优收费应为多少,以最小化平均行程时间并避免Braess悖论?
  • RQ4在真实世界网络中,Boulevard路段的最优信号控制时移为多少,可使可控智能体的平均行程时间最小化?
  • RQ5收费与信号控制如何影响自私智能体的行为及整体系统效率?

主要发现

  • 通过调整奖励函数,MARL范式将DUE与DSO推广至无模型、数据驱动的场景。
  • 在Braess网络中,中央路段收费至少为25时,可最小化平均行程时间并避免Braess悖论。
  • 在真实世界的69节点、166条边网络中,Boulevard路段的最优信号时移为4秒,可使可控智能体的平均行程时间最小化。
  • 当奖励反映个体成本最小时,MARL系统的均衡结果与DUE一致。
  • 当奖励反映系统整体成本最小时,MARL系统的均衡结果与DSO一致。
  • 收费与信号配时等行政调控手段能有效引导智能体行为趋向系统最优结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。