Skip to main content
QUICK REVIEW

[Paper Review] Close the Gaps: A Learning-while-Doing Algorithm for a Class of Single-Product Revenue Management Problems

Zizhuo Wang, Shiming Deng|arXiv (Cornell University)|Jan 24, 2011
Advanced Bandit Algorithms ResearchDecision Sciences21 references16 citations
TL;DR

This paper proposes a dynamic 'learning-while-doing' algorithm for single-product revenue management under demand uncertainty, where the retailer learns the optimal price in real time by iteratively testing prices within shrinking intervals. The method achieves an asymptotic regret of $ O^*(n^{-1/2}) $, which is among the fastest possible rates, outperforming both nonparametric and parametric approaches, especially under model misspecification.

ABSTRACT

We consider a retailer selling a single product with limited on-hand inventory over a finite selling season. Customer demand arrives according to a Poisson process, the rate of which is influenced by a single action taken by the retailer (such as price adjustment, sales commission, advertisement intensity, etc.). The relationship between the action and the demand rate is not known in advance. However, the retailer is able to learn the optimal action "on the fly" as she maximizes her total expected revenue based on the observed demand reactions. Using the pricing problem as an example, we propose a dynamic "learning-while-doing" algorithm that only involves function value estimation to achieve a near-optimal performance. Our algorithm employs a series of shrinking price intervals and iteratively tests prices within that interval using a set of carefully chosen parameters. We prove that the convergence rate of our algorithm is among the fastest of all possible algorithms in terms of asymptotic "regret" (the relative loss comparing to the full information optimal solution). Our result closes the performance gaps between parametric and non-parametric learning and between a post-price mechanism and a customer-bidding mechanism. Important managerial insight from this research is that the values of information on both the parametric form of the demand function as well as each customer's exact reservation price are less important than prior literature suggests. Our results also suggest that firms would be better off to perform dynamic learning and action concurrently rather than sequentially.

Motivation & Objective

  • To address the challenge of revenue management when the demand function is unknown and must be learned in real time.
  • To close the performance gap between parametric and nonparametric learning approaches in revenue management.
  • To demonstrate that concurrent learning and action outperform sequential exploration-exploitation strategies.
  • To quantify the value of information on parametric demand forms and individual customer reservation prices.
  • To develop a robust, nonparametric pricing algorithm that maintains high performance across diverse demand function families.

Proposed method

  • The algorithm uses a series of shrinking price intervals that always contain the optimal price with high probability.
  • It performs iterative price experimentation within each interval using carefully chosen parameters to balance exploration and exploitation.
  • Function value estimation is the only required operation, avoiding complex derivative or model fitting.
  • The algorithm maintains a confidence interval around the optimal price and progressively narrows it based on observed demand responses.
  • A worst-case regret bound is derived to prove asymptotic optimality, showing the algorithm's performance is near-the-best possible.
  • The method is extended to handle kinked demand functions by incorporating prior knowledge of structural breaks.

Experimental results

Research questions

  • RQ1What is the optimal learning strategy for nonparametric revenue management when the demand function is unknown?
  • RQ2How does the regret of a dynamic learning algorithm compare to one-time grid-based learning in terms of convergence rate?
  • RQ3How valuable is prior knowledge of the parametric form of the demand function in a nonparametric setting?
  • RQ4What is the relative benefit of knowing individual customer reservation prices versus only observing purchase decisions?
  • RQ5Can learning and action be effectively combined in a single dynamic procedure to achieve superior performance?

Key findings

  • The proposed dynamic pricing algorithm (DPA) achieves an asymptotic regret of $ O^*(n^{-1/2}) $, which is among the fastest possible rates across all algorithms.
  • In simulations, DPA consistently outperformed the nonparametric policy from Besbes and Zeevi (2013) across all tested values of $ n $, with significantly lower regret.
  • Parametric policies (P-BZ-L and P-BZ-E) performed well only when the true demand function matched their assumed form, but failed to converge to zero regret under misspecification.
  • The regret of DPA decreased with a slope of approximately -0.5 on a log-log scale, indicating a stable $ n^{-1/2} $ convergence rate.
  • Even with optimal learning point selection, parametric policies underperformed DPA as $ n $ increased, highlighting the risk of model misspecification.
  • The results suggest that firms should prioritize dynamic, concurrent learning and action over sequential learning or relying on parametric assumptions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.