[论文解读] A Survey of Contextual Optimization Methods for Decision Making under Uncertainty
对情境优化的全面综述,详细介绍三个学习与优化框架(决策规则优化、序贯学习与优化、以及综合学习与优化)及其模型、训练方法和理论保证。
Recently there has been a surge of interest in operations research (OR) and the machine learning (ML) community in combining prediction algorithms and optimization techniques to solve decision-making problems in the face of uncertainty. This gave rise to the field of contextual optimization, under which data-driven procedures are developed to prescribe actions to the decision-maker that make the best use of the most recently updated information. A large variety of models and methods have been presented in both OR and ML literature under a variety of names, including data-driven optimization, prescriptive optimization, predictive stochastic programming, policy optimization, (smart) predict/estimate-then-optimize, decision-focused learning, (task-based) end-to-end learning/forecasting/optimization, etc. Focusing on single and two-stage stochastic programming problems, this review article identifies three main frameworks for learning policies from data and discusses their strengths and limitations. We present the existing models and methods under a uniform notation and terminology and classify them according to the three main frameworks identified. Our objective with this survey is to both strengthen the general understanding of this active field of research and stimulate further theoretical and algorithmic advancements in integrating ML and stochastic programming.
研究动机与目标
- 阐明如何使用侧信息(协变量)在不确定性下为决策提供信息。
- 统一决策规则优化、序列学习与优化,以及综合学习与优化之间的记号与术语。
- 总结文献中的模型、训练过程以及理论保证。
- 强调将机器学习与随机规划相结合的开放问题与方向。
提出的方法
- 用协变量和不确定参数定义情境优化问题。
- 提出三种学习范式:决策规则优化、序列学习与优化(SLO)、以及综合学习与优化(ILO)。
- 在决策规则框架内回顾线性、基于RKHS的以及非线性决策规则。
- 在ILO中讨论分布鲁robust(分布鲁鲁robust)与代理/可微训练方法。
- 通过展开、隐式求导以及可微代理(如 SPO+)来实现训练。
- 总结与策略优化、端到端学习等相关范式的联系。
实验结果
研究问题
- RQ1在情境优化中用于学习策略的主要框架有哪些,它们有何不同?
- RQ2在情境信息下,不同的决策规则(线性、RKHS、非线性)的表现如何?
- RQ3哪种训练范例最好地将预测模型与下游优化目标对齐?
- RQ4关于这些情境优化方法存在哪些理论保证,包括鲁棒性和一致性?
- RQ5在将ML与随机规划结合时,存在哪些尚待解决的理论与算法挑战?
主要发现
- 识别出三大框架:决策规则优化、序列学习与优化(SLO)、以及综合学习与优化(ILO)。
- 基于RKHS的和非线性决策规则可以超越线性策略,在某些设置中实现渐近最优。
- 综合学习强调直接为处方性能优化预测模型,而不仅仅是预测精度。
- 探索分布式鲁棒和以 Wasserstein 为基础的方法,以防止模型错配和数据漂移。
- 该综述将框架与相关工作如后悔最小化和端到端学习联系起来,并讨论通过展开和隐式求导进行训练。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。