Skip to main content
QUICK REVIEW

[论文解读] Bulk-Calibrated Credal Ambiguity Sets: Fast, Tractable Decision Making under Out-of-Sample Contamination

Mengqi Chen, Thomas B. Berrett|arXiv (Cornell University)|Jan 29, 2026
Risk and Portfolio Optimization被引用 0
一句话总结

论文提出基于大块校准的甄别模糊不确定性集(LV),将不精确概率转化为可处理的DRO目标,从而在数据驱动的大块校准下对样本外污染实现快速鲁棒决策。

ABSTRACT

Distributionally robust optimisation (DRO) minimises the worst-case expected loss over an ambiguity set that can capture distributional shifts in out-of-sample environments. While Huber (linear-vacuous) contamination is a classical minimal-assumption model for an $\varepsilon$-fraction of arbitrary perturbations, including it in an ambiguity set can make the worst-case risk infinite and the DRO objective vacuous unless one imposes strong boundedness or support assumptions. We address these challenges by introducing bulk-calibrated credal ambiguity sets: we learn a high-mass bulk set from data while considering contamination inside the bulk and bounding the remaining tail contribution separately. This leads to a closed-form, finite $\mathrm{mean}+\sup$ robust objective and tractable linear or second-order cone programs for common losses and bulk geometries. Through this framework, we highlight and exploit the equivalence between the imprecise probability (IP) notion of upper expectation and the worst-case risk, demonstrating how IP credal sets translate into DRO objectives with interpretable tolerance levels. Experiments on heavy-tailed inventory control, geographically shifted house-price regression, and demographically shifted text classification show competitive robustness-accuracy trade-offs and efficient optimisation times, using Bayesian, frequentist, or empirical reference distributions.

研究动机与目标

  • 在分布不确定性和样本外污染下提升鲁棒决策的动机。
  • 引入大块受限的甄别模糊不确定性集(前向LV),以获得封闭形式的 Worst-Case 风险。
  • 提供具有有限样本保证和高概率风险证书的数据驱动大块校准。
  • 证明 IP 甄别集与可解释容忍度水平的 DRO 目标等价。

提出的方法

  • 围绕数据驱动的中心分布定义一个大块受限的 LV 甄别模糊不确定性集。
  • 推导封闭式 Worst-Case 风险: (1−ε) E_{P_c,Ξ0}[f_x(ξ)] + ε sup_{ξ∈Ξ0} f_x(ξ)。
  • 为常见损失和大块几何形状提供可处理的 LP/SOCP 重表述。
  • 使用基于分数的选择和基于 DKW 的风险证书来校准 Ξ0 的大块集。
  • 证明风险上界在大块鲁棒性与在 Hubber ε-污染下的尾部控制之间进行分离。
  • 证明 IP 上界期望与一个 DRO 最坏情况风险之间的等价性。
Figure 2 : Worst-case distributions $Q^{\star}$ for $\sup_{Q}\mathbb{E}_{\xi\sim Q}[f]$ under forward LV, reverse LV, and TV balls around a centre $\mathbb{P}_{c,\Xi_{0}}$ (loss $f$ plateaus at a small region to avoid Dirac deltas).
Figure 2 : Worst-case distributions $Q^{\star}$ for $\sup_{Q}\mathbb{E}_{\xi\sim Q}[f]$ under forward LV, reverse LV, and TV balls around a centre $\mathbb{P}_{c,\Xi_{0}}$ (loss $f$ plateaus at a small region to avoid Dirac deltas).

实验结果

研究问题

  • RQ1如何在在高斯误差 ε 污染下构造一个在无限空间内仍然良定义的鲁棒优化目标?
  • RQ2能否从数据中学习具有质量保证的大块集合,以实现可处理的DRO 表达式?
  • RQ3不精确概率(甄别集)与持续空间中的分布鲁棒优化之间的关系是什么?
  • RQ4基于 LV 的甄别集是否在现实任务中提供具有竞争力的鲁棒性与性能?
  • RQ5大块校准如何在不同数据集和损失函数下影响计算效率与鲁棒性之间的权衡?

主要发现

  • 大块受限的 LV 甄别模糊不确定性集合给出封闭形式的 Worst-Case 风险: (1−ε) E_{P_c,Ξ0}[f_x(ξ)] + ε sup_{ξ∈Ξ0} f_x(ξ)。
  • 该方法为常见损失和大块几何形状提供了可处理的 LP 或 SOCP 重表述。
  • 使用基于 DKW 的分数选择对 Ξ0 进行数据校准,提供高概率的大块质量证书(1−γ,置信度为 1−δ)。
  • 在重尾库存控制、灾备迁移下的加州房价回归和 CivilComments 文本分类等任务上,鲁棒性-准确度权衡具有竞争力且优化时间比基线更快。
  • 与 KL 基 DRO 和 OR-WDRO 基线相比,LV 基方法在污染情形下通常表现出更优的 OOS 性能和更短的求解时间。
  • 该框架兼容贝叶斯、频率派或经验参考分布,并提供灵活的中心选择。
Figure 3 : Student- $t$ newsvendor (cost: lower-left is better). Top row: OOS mean–variance frontiers for a range of $\varepsilon_{\operatorname{LV}}\in(0,1]$ ; $\varepsilon_{\operatorname{KL}}\in(0,25]$ . OR-WDRO uses $\varepsilon_{\operatorname{LV}}$ in $(0,0.5)$ . Each point represents one $\vare
Figure 3 : Student- $t$ newsvendor (cost: lower-left is better). Top row: OOS mean–variance frontiers for a range of $\varepsilon_{\operatorname{LV}}\in(0,1]$ ; $\varepsilon_{\operatorname{KL}}\in(0,25]$ . OR-WDRO uses $\varepsilon_{\operatorname{LV}}$ in $(0,0.5)$ . Each point represents one $\vare

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。