[Paper Review] Constant Approximation for Network Revenue Management with Markovian-Correlated Customer Arrivals
This paper proposes a Markovian state-based model to capture time-correlated customer arrivals in network revenue management (NRM), where system states evolve via a time-inhomogeneous Markov chain. It introduces a new linear programming (LP) approximation that serves as an asymptotically optimal upper bound on the optimal policy’s expected reward, and develops a bid price policy with a theoretical approximation ratio of $1/(1+L)$, where $L$ is the maximum number of resources a customer may require.
The Network Revenue Management (NRM) problem is a well-known challenge in dynamic decision-making under uncertainty. In this problem, fixed resources must be allocated to serve customers over a finite horizon, while customers arrive according to a stochastic process. The typical NRM model assumes that customer arrivals are independent over time. However, in this paper, we explore a more general setting where customer arrivals over different periods can be correlated. We propose a model that assumes the existence of a system state, which determines customer arrivals for the current period. This system state evolves over time according to a time-inhomogeneous Markov chain. We show our model can be used to represent correlation in various settings. To solve the NRM problem under our correlated model, we derive a new linear programming (LP) approximation of the optimal policy. Our approximation provides an upper bound on the total expected value collected by the optimal policy. We use our LP to develop a new bid price policy, which computes bid prices for each system state and time period in a backward induction manner. The decision is then made by comparing the reward of the customer against the associated bid prices. Our policy guarantees to collect at least $1/(1+L)$ fraction of the total reward collected by the optimal policy, where $L$ denotes the maximum number of resources required by a customer. In summary, our work presents a Markovian model for correlated customer arrivals in the NRM problem and provides a new LP approximation for solving the problem under this model. We derive a new bid price policy and provides a theoretical guarantee of the performance of the policy.
Motivation & Objective
- To model correlated customer arrivals in NRM beyond independent arrival assumptions, which fail to capture high variance and non-stationary demand patterns.
- To address the computational intractability of the optimal policy under correlated arrivals by developing a tractable approximation method.
- To design a near-optimal bid price policy with strong theoretical performance guarantees under the new correlated model.
- To extend the framework to the assortment setting, where customers choose from multiple products based on a choice model.
Proposed method
- Introduces a system state that determines customer arrivals at each period, evolving via a time-inhomogeneous Markov chain to model temporal correlation.
- Develops a new linear programming (LP) relaxation that upper-bounds the optimal policy’s expected reward and is asymptotically optimal as initial capacities scale up.
- Derives a backward-induction-based bid price policy that computes state- and time-specific bid prices for each customer arrival.
- Applies the LP approximation to guide a bid price control policy that accepts a customer only if the reward exceeds the sum of associated bid prices.
- Extends the framework to the assortment setting by modeling customer choice behavior and adapting the bid price policy accordingly.
- Empirically evaluates two algorithms—BBP and ADP heuristics—against the LP upper bound to assess performance.
Experimental results
Research questions
- RQ1How can we model correlated customer arrivals in NRM to capture both high variance and non-stationary demand patterns in a unified framework?
- RQ2Can we design a computationally tractable approximation of the optimal policy under correlated arrivals that provides a strong theoretical performance guarantee?
- RQ3What is the best achievable approximation ratio for a bid price policy under this correlated model, and can it be generalized to the assortment setting?
- RQ4How do the proposed algorithms perform in practice compared to the optimal policy and the LP upper bound?
Key findings
- The proposed LP approximation serves as an asymptotically optimal upper bound on the expected reward of the optimal policy as initial capacities grow large.
- The bid price policy achieves a theoretical approximation ratio of $1/(1+L)$, where $L$ is the maximum number of resources required by any customer.
- Numerical experiments show that the BBP algorithm achieves an average gap of 6.60% from the LP upper bound, while the ADP heuristics achieve a 6.62% gap on average.
- In specific settings, the BBP algorithm’s performance gap to the LP upper bound ranges from 2.67% to 11.52%, with ADP heuristics showing similar or slightly better performance.
- The model successfully generalizes to the assortment setting, where customers choose from multiple products based on a choice model, extending applicability to real-world scenarios.
- The results demonstrate that the proposed policy performs reasonably well in practice, with small gaps relative to the LP upper bound, validating its practical relevance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.