[论文解读] On the Hardness of Inventory Management with Censored Demand Data
本文提出了一种非随机性、最小化遗憾的重复报童问题框架,适用于右删失需求场景,即仅能观测到销量(而非真实需求)。通过采用具有无偏成本估计的指数加权预测器,该方法在时间周期、库存选择和需求支撑集方面实现了最优遗憾尺度(对数因子以内),表明在遗憾准则下,删失与信息稀缺对性能的影响可忽略不计。
We consider a repeated newsvendor problem where the inventory manager has no prior information about the demand, and can access only censored/sales data. In analogy to multi-armed bandit problems, the manager needs to simultaneously "explore" and "exploit" with her inventory decisions, in order to minimize the cumulative cost. We make no probabilistic assumptions---importantly, independence or time stationarity---regarding the mechanism that creates the demand sequence. Our goal is to shed light on the hardness of the problem, and to develop policies that perform well with respect to the regret criterion, that is, the difference between the cumulative cost of a policy and that of the best fixed action/static inventory decision in hindsight, uniformly over all feasible demand sequences. We show that a simple randomized policy, termed the Exponentially Weighted Forecaster, combined with a carefully designed cost estimator, achieves optimal scaling of the expected regret (up to logarithmic factors) with respect to all three key primitives: the number of time periods, the number of inventory decisions available, and the demand support. Through this result, we derive an important insight: the benefit from "information stalking" as well as the cost of censoring are both negligible in this dynamic learning problem, at least with respect to the regret criterion. Furthermore, we modify the proposed policy in order to perform well in terms of the tracking regret, that is, using as benchmark the best sequence of inventory decisions that switches a limited number of times. Numerical experiments suggest that the proposed approach outperforms existing ones (that are tailored to, or facilitated by, time stationarity) on nonstationary demand models. Finally, we extend the proposed approach and its analysis to a "combinatorial" version of the repeated newsvendor problem.
研究动机与目标
- 解决仅能观测到销量(而非真实需求)的库存管理挑战。
- 开发一种在非平稳、对抗性需求序列下表现良好的策略,且无需概率假设。
- 采用遗憾准则评估性能,与事后最优固定库存决策进行比较。
- 将框架扩展至组合设置,如单仓多零售商系统。
- 在跟踪遗憾下分析性能,其中基准为切换次数有限的最佳动作序列。
提出的方法
- 采用非随机(对抗性)框架,将需求建模为任意序列,不施加分布假设。
- 引入指数加权预测器(EWF)以实现动态库存决策,兼顾探索与利用。
- 使用无偏成本估计器从删失销量数据中重构成本差异,从而实现准确的遗憾分析。
- 在缺乏完整反馈的情况下,采用带指数加权的随机化策略以平衡探索与利用。
- 将EWF扩展为跟随扰动线性函数(FPL-IX)变体,以获得跟踪遗憾的保证。
- 通过将估计器和策略泛化至多地点设置,将框架适配至组合报童问题。
实验结果
研究问题
- RQ1当仅能获取删失需求(销量)数据时,库存管理的根本难度是什么?
- RQ2在无概率假设下,能否在对抗性、非平稳需求下实现最优遗憾尺度?
- RQ3删失成本与信息获取需求在遗憾意义上的影响如何?
- RQ4该框架能否扩展至组合库存系统(如多零售商)并保持类似的遗憾保证?
- RQ5在跟踪遗憾下,可实现怎样的性能边界,其中基准允许库存策略有限次切换?
主要发现
- 结合无偏成本估计的指数加权预测器在时间、动作数量和需求支撑集大小方面实现了最优期望遗憾尺度(对数因子以内)。
- 信息追踪的收益与删失成本在遗憾意义上可忽略不计,表明对有限反馈具有鲁棒性。
- 在数值实验中,所提策略在非平稳需求模型下优于针对平稳需求设计的现有方法。
- FPL-IX变体提供了非平凡的跟踪遗憾边界,支持对允许有限次动作切换的基准的性能表现。
- 该框架可扩展至组合报童问题,在多零售商设置中实现了近似最优的遗憾性能。
- 结果在所有可行需求序列上一致成立,无需平稳性、独立性或分布假设。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。