Skip to main content
QUICK REVIEW

[Paper Review] Dynamic Pricing and Management for Electric Autonomous Mobility on Demand Systems Using Reinforcement Learning.

Berkay Turan, Ramtin Pedarsani|arXiv (Cornell University)|Sep 16, 2019
Transportation and Mobility Innovations4 citations
TL;DR

This paper proposes a deep reinforcement learning (DRL)-based dynamic pricing and fleet management policy for electric autonomous mobility-on-demand systems, modeling the system as a Markov decision process to handle time-varying trip demand, energy prices, and renewable availability. The DRL policy achieves up to 200× lower queue lengths and higher profits than static baselines in Manhattan and San Francisco case studies.

ABSTRACT

The proliferation of ride sharing systems is a major drive in the advancement of autonomous and electric vehicle technologies. This paper considers the joint routing, battery charging, and pricing problem faced by a profit-maximizing transportation service provider that operates a fleet of autonomous electric vehicles. We define the dynamic system model that captures the time dependent and stochastic features of an electric autonomous-mobility-on-demand system. To accommodate for the time-varying nature of trip demands, renewable energy availability, and electricity prices and to further optimally manage the autonomous fleet, a dynamic policy is required. In order to develop a dynamic control policy, we first formulate the dynamic progression of the system as a Markov decision process. We argue that it is intractable to exactly solve for the optimal policy using exact dynamic programming methods and therefore apply deep reinforcement learning to develop a near-optimal control policy. Furthermore, we establish the static planning problem by considering time-invariant system parameters. We define the capacity region and determine the optimal static policy to serve as a baseline for comparison with our dynamic policy. While the static policy provides important insights on optimal pricing and fleet management, we show that in a real dynamic setting, it is inefficient to utilize a static policy. The two case studies we conducted in Manhattan and San Francisco demonstrate the efficacy of our dynamic policy in terms of network stability and profits, while keeping the queue lengths up to 200 times less than the static policy.

Motivation & Objective

  • To address the joint challenge of routing, battery charging, and dynamic pricing in profit-maximizing electric autonomous mobility-on-demand systems.
  • To model the time-varying and stochastic nature of trip demand, electricity prices, and renewable energy availability in a dynamic system framework.
  • To develop a near-optimal control policy using deep reinforcement learning due to the intractability of exact dynamic programming solutions.
  • To establish a static benchmark policy via a time-invariant planning problem to compare performance against the dynamic policy.
  • To evaluate the efficacy of the dynamic policy in real-world urban settings, focusing on network stability and profitability.

Proposed method

  • Formulate the electric autonomous mobility-on-demand system as a Markov decision process (MDP) to capture time-dependent and stochastic dynamics.
  • Apply deep reinforcement learning (DRL) to learn a near-optimal control policy due to the intractability of exact dynamic programming for large-scale systems.
  • Define a static planning problem with time-invariant parameters to derive a capacity region and optimal static policy as a performance baseline.
  • Use case studies in Manhattan and San Francisco to simulate real-world conditions, including fluctuating demand, energy prices, and renewable availability.
  • Train the DRL agent to optimize long-term profit by balancing routing, charging decisions, and dynamic pricing in response to system state changes.
  • Compare the dynamic DRL policy against the static policy in terms of queue lengths, system stability, and profit generation.

Experimental results

Research questions

  • RQ1How does a dynamic pricing and fleet management policy based on deep reinforcement learning compare to a static policy in terms of system stability and profitability in electric autonomous mobility-on-demand systems?
  • RQ2To what extent can a DRL-based policy reduce queue lengths in high-demand urban environments like Manhattan and San Francisco?
  • RQ3What is the impact of time-varying factors such as trip demand, electricity prices, and renewable energy availability on fleet management and pricing decisions?
  • RQ4How does the dynamic policy perform relative to the static policy in maintaining network stability under stochastic and non-stationary conditions?
  • RQ5Can a learned DRL policy effectively balance routing, charging, and pricing to maximize long-term profit in electric autonomous vehicle systems?

Key findings

  • The dynamic DRL policy reduces queue lengths by up to 200 times compared to the static policy in both Manhattan and San Francisco case studies.
  • The dynamic policy achieves higher profits than the static policy while maintaining significantly better network stability under time-varying demand and energy conditions.
  • The static policy, while useful as a baseline, is inefficient in real-world dynamic environments due to its inability to adapt to changing system parameters.
  • The DRL-based approach effectively balances routing, battery charging, and dynamic pricing in response to real-time fluctuations in demand and electricity prices.
  • The capacity region derived from the static planning problem provides a theoretical foundation for evaluating the performance of dynamic policies.
  • The case studies confirm that dynamic control policies are essential for scalable and profitable operation of electric autonomous mobility-on-demand systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.