[论文解读] Robust Dynamic Assortment Optimization in the Presence of Outlier Customers
本文提出了一种针对多项对数(MNL)模型在 $\varepsilon$-污染环境下的鲁棒在线组合优化策略,其中部分比例为 $\varepsilon$ 的客户为异常值,会做出任意选择。通过采用主动排除策略,该方法实现了近乎最优的遗憾边界——当组合容量为常数时,遗憾随 $T$ 对数增长——且在未知 $\varepsilon$ 的情况下依然有效,模拟结果表明其优于非鲁棒的UCB和Thompson Sampling方法。
We consider the dynamic assortment optimization problem under the multinomial logit model (MNL) with unknown utility parameters. The main question investigated in this paper is model mis-specification under the $\varepsilon$-contamination model, which is a fundamental model in robust statistics and machine learning. In particular, throughout a selling horizon of length $T$, we assume that customers make purchases according to a well specified underlying multinomial logit choice model in a $(1-\varepsilon)$-fraction of the time periods, and make arbitrary purchasing decisions instead in the remaining $\varepsilon$-fraction of the time periods. In this model, we develop a new robust online assortment optimization policy via an active elimination strategy. We establish both upper and lower bounds on the regret, and show that our policy is optimal up to logarithmic factor in $T$ when the assortment capacity is constant. %% capacity of assortments has a constant upper limit. We further develop a fully adaptive policy that does not require any prior knowledge of the contamination parameter $\varepsilon$. In the case of the existence a sub-optimality gap between optimal and sub-optimal products, we also established gap-dependent logarithmic regret upper bounds and lower bounds in both the known-$\varepsilon$ and unknown-$\varepsilon$ cases. Our simulation study shows that our policy outperforms the existing policies based on upper confidence bounds (UCB) and Thompson sampling.
研究动机与目标
- 解决由于偏离标准MNL选择模型的异常客户导致的动态组合优化中的模型误设问题。
- 在在线学习环境中设计一种对对抗性污染具有鲁棒性的决策策略,其中部分比例为 $\varepsilon$ 的客户会做出任意购买行为。
- 在已知与未知 $\varepsilon$ 的情况下建立紧致的遗憾边界,包括基于差距的分析。
- 设计一种无需事先知晓污染水平 $\varepsilon$ 的自适应策略。
- 通过实证结果证明,所提方法在存在污染的现实场景下,优于非鲁棒基线方法如UCB和Thompson Sampling。
提出的方法
- 本文提出一种基于新型主动排除策略的政策,该策略在MNL模型下,利用鲁棒置信区间系统性地排除次优产品。
- 将客户行为建模为一个良好设定的MNL分布(用于典型客户)与一个任意污染分布(用于异常值)的混合,污染水平为 $\varepsilon$。
- 该策略采用鲁棒估计框架,计算对异常值具有鲁棒性的置信区间边界,即使污染分布 $Q_t$ 随时间变化也保持稳健。
- 在未知 $\varepsilon$ 的情况下,该方法采用完全自适应策略,基于观测数据动态调整,无需事先知晓 $\varepsilon$,并结合倍增技巧与置信区间收紧。
- 遗憾分析利用基于差距的边界,其中最优与次优商品之间的次优差距影响收敛速度。
- 通过遗憾的上下界分析建立了理论保证,表明当组合容量为常数时,遗憾在 $T$ 的对数因子范围内近乎最优。
实验结果
研究问题
- RQ1我们能否设计一种动态组合优化策略,使得当 $\varepsilon$ 比例的客户为异常值并做出任意选择(而非遵循假设的MNL模型)时,策略依然有效?
- RQ2在在线组合优化中存在此类对抗性污染的情况下,遗憾的根本极限(下界)是什么?
- RQ3鲁棒策略的性能如何依赖于最优与次优产品之间的次优差距?
- RQ4我们能否开发一种自适应策略,实现在未知 $\varepsilon$ 情况下的良好遗憾性能?
- RQ5在现实的异常值场景下,所提出的鲁棒策略与非鲁棒方法(如UCB和Thompson Sampling)相比,其表现如何?
主要发现
- 在已知 $\varepsilon$ 的设定下,所提出的主动排除策略在组合容量为常数时,遗憾边界在 $T$ 的对数因子范围内达到最优。
- 在基于差距的范式中,该策略实现了随次优差距缩放的对数遗憾上界,证实当最优产品显著优于次优产品时收敛速度更快。
- 在未知 $\varepsilon$ 的情况下,自适应策略实现了与已知 $\varepsilon$ 策略相同阶的遗憾,表明其在无需污染知识的情况下仍具鲁棒性。
- 模拟结果表明,当 $\varepsilon > 0$ 时,所提方法的平均遗憾稳定在极低水平(0.02–0.06),远低于UCB和Thompson Sampling,且其遗憾随时间跨度 $T$ 增大而减小,而基线方法则无此趋势。
- 当 $\varepsilon = 0$ 时,所提方法性能略逊于基线方法,但其平均遗憾的下降速率保持一致,表明在无污染情况下仅带来可忽略的额外开销。
- 该方法对随时间变化的异常值分布 $Q_t$ 具有鲁棒性,使其比静态污染模型更具实际应用价值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。