[Paper Review] Bulk-Calibrated Credal Ambiguity Sets: Fast, Tractable Decision Making under Out-of-Sample Contamination
The paper introduces bulk-calibrated credal ambiguity sets (LV) that translate imprecise probability into a tractable DRO objective, enabling fast robust decisions under out-of-sample contamination with data-driven bulk calibration.
Distributionally robust optimisation (DRO) minimises the worst-case expected loss over an ambiguity set that can capture distributional shifts in out-of-sample environments. While Huber (linear-vacuous) contamination is a classical minimal-assumption model for an $\varepsilon$-fraction of arbitrary perturbations, including it in an ambiguity set can make the worst-case risk infinite and the DRO objective vacuous unless one imposes strong boundedness or support assumptions. We address these challenges by introducing bulk-calibrated credal ambiguity sets: we learn a high-mass bulk set from data while considering contamination inside the bulk and bounding the remaining tail contribution separately. This leads to a closed-form, finite $\mathrm{mean}+\sup$ robust objective and tractable linear or second-order cone programs for common losses and bulk geometries. Through this framework, we highlight and exploit the equivalence between the imprecise probability (IP) notion of upper expectation and the worst-case risk, demonstrating how IP credal sets translate into DRO objectives with interpretable tolerance levels. Experiments on heavy-tailed inventory control, geographically shifted house-price regression, and demographically shifted text classification show competitive robustness-accuracy trade-offs and efficient optimisation times, using Bayesian, frequentist, or empirical reference distributions.
Motivation & Objective
- Motivate robust decision-making under distributional uncertainty and out-of-sample contamination.
- Introduce bulk-restricted credal ambiguity sets (forward LV) that yield a closed-form worst-case risk.
- Provide data-driven bulk calibration with finite-sample guarantees and a high-probability risk certificate.
- Show that IP credal sets correspond to DRO objectives with interpretable tolerance levels.
Proposed method
- Define a bulk-restricted LV credal ambiguity set around a data-driven centre distribution.
- Derive the closed-form worst-case risk: (1−ε) E_{P_c,Ξ0}[f_x(ξ)] + ε sup_{ξ∈Ξ0} f_x(ξ).
- Provide tractable LP/SOCP reformulations for common losses and bulk geometries.
- Calibrate bulk set Ξ0 using score-based selection with DKW-based risk certificates.
- Prove risk bounds that separate in-bulk robustness from tail control under Huber ε-contamination.
- Demonstrate equivalence between IP upper expectation and a DRO worst-case risk.
![Figure 2 : Worst-case distributions $Q^{\star}$ for $\sup_{Q}\mathbb{E}_{\xi\sim Q}[f]$ under forward LV, reverse LV, and TV balls around a centre $\mathbb{P}_{c,\Xi_{0}}$ (loss $f$ plateaus at a small region to avoid Dirac deltas).](https://ar5iv.labs.arxiv.org/html/2601.21324/assets/x1.png)
Experimental results
Research questions
- RQ1How can we construct a robust optimization objective that remains well-posed in unbounded spaces under Huber ε-contamination?
- RQ2Can we learn a bulk set with mass guarantees from data to enable tractable DRO formulations?
- RQ3What is the relationship between imprecise probabilities (credal sets) and distributionally robust optimization in continuous spaces?
- RQ4Do LV-based credal sets provide practical robustness with competitive performance across real-world tasks?
- RQ5How does bulk calibration affect computational efficiency and robustness trade-offs across datasets and losses?
Key findings
- The bulk-restricted LV credal ambiguity set yields a closed-form worst-case risk: (1−ε) E_{P_c,Ξ0}[f_x(ξ)] + ε sup_{ξ∈Ξ0} f_x(ξ).
- The approach leads to tractable LP or SOCP reformulations for common losses and bulk geometries.
- Calibrating Ξ0 with data using a DKW-based score selection provides a high-probability bulk-mass certificate (1−γ with confidence 1−δ).
- Experiments on heavy-tailed inventory control, California housing regression under deployment shift, and CivilComments text classification show competitive robustness-accuracy trade-offs and faster optimisation times than baselines.
- LV-based methods often achieve superior OOS performance and shorter solve times, particularly under contamination, compared with KL-based DRO and OR-WDRO baselines.
- The framework accommodates Bayesian, frequentist, or empirical reference distributions and offers flexible centre choices.
![Figure 3 : Student- $t$ newsvendor (cost: lower-left is better). Top row: OOS mean–variance frontiers for a range of $\varepsilon_{\operatorname{LV}}\in(0,1]$ ; $\varepsilon_{\operatorname{KL}}\in(0,25]$ . OR-WDRO uses $\varepsilon_{\operatorname{LV}}$ in $(0,0.5)$ . Each point represents one $\vare](https://ar5iv.labs.arxiv.org/html/2601.21324/assets/x2.png)
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.