[论文解读] Online Stochastic Optimization with Wasserstein Based Non-stationarity
本文提出一种基于Wasserstein距离的非平稳性度量方法,用于建模在线随机优化中具有多重预算约束的分布漂移问题。该研究提出了一种信息性梯度下降(IGDP)算法,通过在对偶空间更新中整合先验分布估计,实现了在非平稳环境下的数据驱动与无信息设定下最优阶的遗憾值。
We consider a general online stochastic optimization problem with multiple budget constraints over a horizon of finite time periods. In each time period, a reward function and multiple cost functions are revealed, and the decision maker needs to specify an action from a convex and compact action set to collect the reward and consume the budget. Each cost function corresponds to the consumption of one budget. In each period, the reward and cost functions are drawn from an unknown distribution, which is non-stationary across time. The objective of the decision maker is to maximize the cumulative reward subject to the budget constraints. This formulation captures a wide range of applications including online linear programming and network revenue management, among others. In this paper, we consider two settings: (i) a data-driven setting where the true distribution is unknown but a prior estimate (possibly inaccurate) is available; (ii) an uninformative setting where the true distribution is completely unknown. We propose a unified Wasserstein-distance based measure to quantify the inaccuracy of the prior estimate in setting (i) and the non-stationarity of the system in setting (ii). We show that the proposed measure leads to a necessary and sufficient condition for the attainability of a sublinear regret in both settings. For setting (i), we propose a new algorithm, which takes a primal-dual perspective and integrates the prior information of the underlying distributions into an online gradient descent procedure in the dual space. The algorithm also naturally extends to the uninformative setting (ii). Under both settings, we show the corresponding algorithm achieves a regret of optimal order. In numerical experiments, we demonstrate how the proposed algorithms can be naturally integrated with the re-solving technique to further boost the empirical performance.
研究动机与目标
- 解决在非平稳、未知分布条件下具有多重预算约束的在线随机优化问题。
- 设计一种统一的基于Wasserstein距离的度量方法,用于量化先验估计的不准确性与环境的非平稳性。
- 设计一种算法,在数据驱动与无信息设定下均实现次线性遗憾值。
- 通过处理未知且非平稳的分布,弥合在线线性规划(平稳)与网络收益管理(已知非i.i.d.)之间的差距。
- 通过在网络收益管理问题上的数值实验,验证所提算法的鲁棒性与性能表现。
提出的方法
- 提出基于Wasserstein距离的偏差预算(WBDB),用于在先验估计或真实分布周围定义不确定性集,以捕捉非平稳性与估计误差。
- 设计信息性梯度下降(IGDP)算法,该算法在对偶空间中执行在线梯度下降,并通过对偶变量更新整合先验分布知识。
- 采用一种重优化启发式策略,定期重新求解上界问题以优化对偶变量,从而在不牺牲计算效率的前提下提升性能。
- 应用原始-对偶框架,基于不确定性集与先验估计所导出的对偶价格做出决策。
- 采用周期性时间重优化策略,以自适应方式校正对偶估计,尤其在非平稳环境中表现优异。
- 将WBDB不确定性集整合进遗憾值分析,推导出实现次线性遗憾值的必要与充分条件。

实验结果
研究问题
- RQ1基于Wasserstein距离的度量能否在在线随机优化中对非平稳性与估计误差提供紧致表征?
- RQ2在非平稳分布下,将先验分布估计整合进对偶空间梯度下降算法是否能实现最优遗憾值?
- RQ3在非平稳设定下,结合在线梯度下降与周期性重优化时,计算效率与性能之间的最优权衡是什么?
- RQ4在真实分布未知且非平稳的条件下,次线性遗憾值在何种条件下可被实现?
- RQ5IGDP算法在不同非平稳强度与估计误差水平下的性能相较于基线方法表现如何?
主要发现
- 所提出的基于Wasserstein距离的偏差预算(WBDB)在数据驱动与无信息设定下,均提供了实现次线性遗憾值的必要与充分条件。
- IGDP算法通过自适应地将先验分布知识整合进对偶空间更新,实现了最优阶的遗憾值。
- 数值实验表明,IGDP在不同非平稳强度(α)与估计误差(β)下均保持稳定性能,其表现达到上界值的94.9%至98.8%。
- 重优化启发式策略显著提升了性能,尤其在低频次(如每100或200个周期)时效果明显,而频率超过100后增益趋于平缓。
- 即使在高估计误差(β = 0.04)条件下,采用重优化的IGDP仍可达到上界值的98.1%至98.8%,展现出强鲁棒性。
- 在非平稳环境中,该算法通过利用先验信息与自适应重优化,显著优于标准OGD类方法。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。