[Paper Review] A comparison of control strategies applied to a pricing problem in retail
This paper compares dynamic pricing control strategies—specifically Certainty Equivalent Control (CEC), Open-Loop Feedback Control (OLFC), and the optimal Bellman policy—in a stochastic retail pricing problem. Using simulation on a one-product model, it finds that CEC outperforms the optimal policy in over 50% of realizations, despite lower average profit, indicating a risk-seeking bias; OLFC provides a strong compromise, closely approximating the Bellman policy with better performance than CEC.
When sales of a product are affected by randomness in demand, retailers can use dynamic pricing strategies to maximise their profits. In this article the pricing problem is formulated as a stochastic optimal control problem, where the optimal policy can be found by solving the associated Bellman equation. The aim is to investigate Approximate Dynamic Programming algorithms for this problem. For realistic retail applications, modelling the problem and solving it to optimality is intractable. Thus practitioners make simplifying assumptions and design suboptimal policies, but a thorough investigation of the relative performance of these policies is lacking. To better understand such assumptions, we simulate the performance of two algorithms on a one-product system. It is found that for more than half of the realisations of the random disturbance, the often-used, but approximate, Certainty Equivalent Control policy yields larger profits than an optimal, maximum expected-value policy. This approximate algorithm, however, performs significantly worse in the remaining realisations, which colloquially can be interpreted as a more risk-seeking attitude by the retailer. Another policy, Open-Loop Feedback Control, is shown to work well as a compromise between the Certainty Equivalent Control and the optimal policy.
Motivation & Objective
- To evaluate the performance of suboptimal control policies—CEC and OLFC—relative to the optimal Bellman policy in a stochastic retail pricing problem.
- To investigate how suboptimal policies affect the full distribution of profits, not just expected values.
- To assess whether practical approximations like CEC and OLFC offer viable alternatives to computationally intractable optimal solutions.
- To highlight the risk implications of choosing approximate policies, especially in terms of profit distribution skewness.
- To provide empirical evidence on policy performance under realistic, randomly disturbed demand.
Proposed method
- Formulates the pricing problem as a stochastic optimal control problem with discrete time, state-dependent stock levels, and random demand disturbances.
- Solves the optimal policy via the Bellman equation, which is computationally intractable for realistic systems.
- Implements the Certainty Equivalent Control (CEC) policy by replacing random disturbances with their expected values in the control law.
- Develops the Open-Loop Feedback Control (OLFC) policy by approximating the expectation of future costs over a horizon, improving on CEC.
- Uses Monte Carlo simulation with 10,000 samples of random disturbances to compare profit distributions across policies.
- Employs statistical analysis and empirical distributions to compare the relative performance of CEC, OLFC, and Bellman policies.
Experimental results
Research questions
- RQ1Does the Certainty Equivalent Control (CEC) policy, despite being suboptimal, yield higher profits than the optimal Bellman policy in more than half of the simulated realizations?
- RQ2How does the profit distribution of the CEC policy compare to that of the optimal Bellman policy, particularly in terms of tail behavior and risk exposure?
- RQ3To what extent does the Open-Loop Feedback Control (OLFC) policy approximate the performance of the optimal Bellman policy in terms of profit distribution and expected value?
- RQ4What is the trade-off between computational cost and performance accuracy when using OLFC versus CEC in dynamic pricing?
- RQ5How do different system parameters (e.g., demand volatility, cost structure) affect the relative performance of the three control strategies?
Key findings
- For more than half of the simulated realizations, the Certainty Equivalent Control (CEC) policy yields higher profits than the optimal Bellman policy, despite the Bellman policy having a higher expected profit.
- The CEC policy generates a larger, lower-tail profit distribution compared to the Bellman policy, indicating a risk-seeking behavior profile in practical terms.
- The Open-Loop Feedback Control (OLFC) policy performs significantly better than CEC, with profit differences an order of magnitude smaller than those between CEC and Bellman.
- The empirical distribution of profit differences between OLFC and Bellman policies is unimodal and concentrated around zero, suggesting OLFC closely approximates the optimal policy.
- The CEC policy is a special case of OLFC with a zeroth-order approximation of future expectations, making OLFC more accurate but computationally more expensive.
- The relative profit difference between OLFC and Bellman policies is consistently an order of magnitude smaller than between CEC and Bellman, confirming OLFC as a strong practical compromise.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.