[论文解读] A comparison of control strategies applied to a pricing problem in retail
本文在随机零售定价问题中比较了动态定价控制策略——具体为确定性等价控制(CEC)、开环反馈控制(OLFC)和最优贝尔曼策略。通过在一产品模型上的仿真,发现尽管平均利润较低,CEC在超过50%的实现中表现优于最优策略,表明存在风险寻求偏差;OLFC则提供了出色的折中方案,其性能在很大程度上接近贝尔曼策略,且优于CEC。
When sales of a product are affected by randomness in demand, retailers can use dynamic pricing strategies to maximise their profits. In this article the pricing problem is formulated as a stochastic optimal control problem, where the optimal policy can be found by solving the associated Bellman equation. The aim is to investigate Approximate Dynamic Programming algorithms for this problem. For realistic retail applications, modelling the problem and solving it to optimality is intractable. Thus practitioners make simplifying assumptions and design suboptimal policies, but a thorough investigation of the relative performance of these policies is lacking. To better understand such assumptions, we simulate the performance of two algorithms on a one-product system. It is found that for more than half of the realisations of the random disturbance, the often-used, but approximate, Certainty Equivalent Control policy yields larger profits than an optimal, maximum expected-value policy. This approximate algorithm, however, performs significantly worse in the remaining realisations, which colloquially can be interpreted as a more risk-seeking attitude by the retailer. Another policy, Open-Loop Feedback Control, is shown to work well as a compromise between the Certainty Equivalent Control and the optimal policy.
研究动机与目标
- 评估次优控制策略——CEC与OLFC——在随机零售定价问题中相对于最优贝尔曼策略的性能。
- 研究次优策略如何影响利润的完整分布,而不仅限于期望值。
- 评估诸如CEC与OLFC等实用近似方法是否可作为计算上不可行的最优解的可行替代方案。
- 强调选择近似策略所带来的风险影响,特别是对利润分布偏度的影响。
- 提供在现实、随机扰动需求下的策略性能实证证据。
提出的方法
- 将定价问题建模为离散时间、状态相关库存水平与随机需求扰动的随机最优控制问题。
- 通过贝尔曼方程求解最优策略,但该方法在实际系统中计算上不可行。
- 通过在控制律中用其期望值替代随机扰动,实现确定性等价控制(CEC)策略。
- 通过在时间范围内近似未来成本的期望值,开发开环反馈控制(OLFC)策略,从而改进CEC。
- 使用10,000次随机扰动样本的蒙特卡洛仿真,比较各策略的利润分布。
- 采用统计分析与经验分布,比较CEC、OLFC与贝尔曼策略的相对性能。
实验结果
研究问题
- RQ1尽管次优,确定性等价控制(CEC)策略是否在超过一半的模拟实现中产生高于最优贝尔曼策略的利润?
- RQ2与最优贝尔曼策略相比,CEC策略的利润分布在尾部分布行为和风险暴露方面有何差异?
- RQ3开环反馈控制(OLFC)策略在利润分布与期望值方面,多大程度上近似于最优贝尔曼策略的性能?
- RQ4在动态定价中,OLFC与CEC相比,计算成本与性能准确度之间的权衡如何?
- RQ5不同系统参数(如需求波动性、成本结构)如何影响三种控制策略的相对性能?
主要发现
- 在超过一半的模拟实现中,尽管贝尔曼策略的期望利润更高,确定性等价控制(CEC)策略仍产生更高的实际利润。
- 与贝尔曼策略相比,CEC策略生成了更大且尾部更低的利润分布,表明其在实际中表现出风险寻求行为特征。
- 开环反馈控制(OLFC)策略显著优于CEC,OLFC与贝尔曼策略之间的利润差异比CEC与贝尔曼策略之间的差异小一个数量级。
- OLFC与贝尔曼策略之间利润差异的经验分布呈单峰且集中在零附近,表明OLFC能很好地近似最优策略。
- CEC策略是OLFC策略的一个特例,其对未来期望的近似为零阶近似,因此OLFC更精确但计算成本更高。
- OLFC与贝尔曼策略之间的相对利润差异始终比CEC与贝尔曼策略之间的差异小一个数量级,证实OLFC是一种出色的实用折中方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。