Skip to main content
QUICK REVIEW

[论文解读] Dynamic Pricing and Management for Electric Autonomous Mobility on Demand Systems Using Reinforcement Learning.

Berkay Turan, Ramtin Pedarsani|arXiv (Cornell University)|Sep 16, 2019
Transportation and Mobility Innovations被引用 4
一句话总结

本文提出了一种基于深度强化学习(DRL)的电动自动驾驶按需出行系统动态定价与车队管理策略,将系统建模为马尔可夫决策过程,以应对随时间变化的出行需求、电价和可再生能源可用性。在曼哈顿和旧金山的案例研究中,该DRL策略相比静态基线策略,将队列长度降低至原来的1/200以下,同时实现更高利润。

ABSTRACT

The proliferation of ride sharing systems is a major drive in the advancement of autonomous and electric vehicle technologies. This paper considers the joint routing, battery charging, and pricing problem faced by a profit-maximizing transportation service provider that operates a fleet of autonomous electric vehicles. We define the dynamic system model that captures the time dependent and stochastic features of an electric autonomous-mobility-on-demand system. To accommodate for the time-varying nature of trip demands, renewable energy availability, and electricity prices and to further optimally manage the autonomous fleet, a dynamic policy is required. In order to develop a dynamic control policy, we first formulate the dynamic progression of the system as a Markov decision process. We argue that it is intractable to exactly solve for the optimal policy using exact dynamic programming methods and therefore apply deep reinforcement learning to develop a near-optimal control policy. Furthermore, we establish the static planning problem by considering time-invariant system parameters. We define the capacity region and determine the optimal static policy to serve as a baseline for comparison with our dynamic policy. While the static policy provides important insights on optimal pricing and fleet management, we show that in a real dynamic setting, it is inefficient to utilize a static policy. The two case studies we conducted in Manhattan and San Francisco demonstrate the efficacy of our dynamic policy in terms of network stability and profits, while keeping the queue lengths up to 200 times less than the static policy.

研究动机与目标

  • 解决在利润最大化的电动自动驾驶按需出行系统中,路径规划、电池充电与动态定价的联合挑战。
  • 在动态系统框架中,对出行需求、电价和可再生能源可用性的时变与随机特性进行建模。
  • 由于精确动态规划解法在大规模系统中难以求解,采用深度强化学习(DRL)开发近似最优控制策略。
  • 通过时间不变的规划问题建立静态基线策略,用于与动态策略进行性能对比。
  • 在真实城市环境中评估动态策略的有效性,重点关注网络稳定性和盈利能力。

提出的方法

  • 将电动自动驾驶按需出行系统建模为马尔可夫决策过程(MDP),以捕捉依赖时间且具有随机性的动态行为。
  • 由于大规模系统中精确动态规划难以求解,采用深度强化学习(DRL)学习近似最优控制策略。
  • 定义一个具有时间不变参数的静态规划问题,以推导容量区域和最优静态策略,作为性能基线。
  • 通过曼哈顿和旧金山的案例研究模拟真实世界条件,包括波动的出行需求、电价和可再生能源可用性。
  • 训练DRL智能体,通过平衡路径选择、充电决策与动态定价,以响应系统状态变化,从而优化长期利润。
  • 从队列长度、系统稳定性与利润生成三个方面,将动态DRL策略与静态策略进行对比。

实验结果

研究问题

  • RQ1在电动自动驾驶按需出行系统中,基于深度强化学习的动态定价与车队管理策略相较于静态策略,在系统稳定性和盈利能力方面表现如何?
  • RQ2DRL策略在曼哈顿和旧金山等高需求城市环境中,能在多大程度上减少队列长度?
  • RQ3出行需求、电价和可再生能源可用性等时变因素对车队管理与定价决策有何影响?
  • RQ4在随机且非平稳条件下,动态策略相较于静态策略在维持网络稳定性方面表现如何?
  • RQ5学习得到的DRL策略能否有效平衡路径规划、充电与定价,以在电动自动驾驶车辆系统中实现长期利润最大化?

主要发现

  • 在曼哈顿和旧金山的案例研究中,与静态策略相比,动态DRL策略将队列长度降低了最多200倍。
  • 在需求和能源条件随时间变化的情况下,该动态策略在实现更高利润的同时,显著提升了网络稳定性。
  • 尽管静态策略可作为基线参考,但在真实世界动态环境中效率低下,因其无法适应系统参数的变化。
  • DRL方法能有效响应需求和电价的实时波动,实现路径规划、电池充电与动态定价的协同平衡。
  • 从静态规划问题推导出的容量区域为评估动态策略的性能提供了理论基础。
  • 案例研究证实,动态控制策略对于实现电动自动驾驶按需出行系统的可扩展与盈利运营至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。