Skip to main content
QUICK REVIEW

[Paper Review] A General Approach to Multi-Armed Bandits Under Risk Criteria

Asaf Cassel, Shie Mannor|arXiv (Cornell University)|Jun 4, 2018
Advanced Bandit Algorithms Research37 citations
TL;DR

This paper presents a unified framework for multi-armed bandits under various risk criteria by identifying general conditions linking reward distributions to performance metrics. It enables systematic design of UCB policies and oracle characterization for objectives like CVaR, mean-variance, and Sharpe ratio, revealing non-intuitive differences in statistical behavior across criteria.

ABSTRACT

Different risk-related criteria have received recent interest in learning problems, where typically each case is treated in a customized manner. In this paper we provide a more systematic approach to analyzing such risk criteria within a stochastic multi-armed bandit (MAB) formulation. We identify a set of general conditions that yield a simple characterization of the oracle rule (which serves as the regret benchmark), and facilitate the design of upper confidence bound (UCB) learning policies. The conditions are derived from problem primitives, primarily focusing on the relation between the arm reward distributions and the (risk criteria) performance metric. Among other things, the work highlights some (possibly non-intuitive) subtleties that differentiate various criteria in conjunction with statistical properties of the arms. Our main findings are illustrated on several widely used objectives such as conditional value-at-risk, mean-variance, Sharpe-ratio, and more.

Motivation & Objective

  • To address the lack of systematic treatment of risk criteria in stochastic multi-armed bandits.
  • To identify general conditions that link arm reward distributions to risk-based performance metrics.
  • To enable a unified characterization of the oracle rule across diverse risk criteria.
  • To facilitate the design of upper confidence bound (UCB) learning policies under these criteria.
  • To highlight subtle statistical differences between risk criteria in bandit settings.

Proposed method

  • Derives general conditions on reward distributions and risk metrics that ensure tractable oracle characterization.
  • Uses problem primitives—specifically, the interplay between reward distribution properties and risk functionals—to define the conditions.
  • Applies the conditions to derive UCB-style learning policies that balance exploration and risk-aware exploitation.
  • Characterizes the optimal policy (oracle) in terms of the risk metric and distributional properties of each arm.
  • Applies the framework to concrete objectives such as conditional value-at-risk, mean-variance, and Sharpe ratio.
  • Demonstrates that different risk criteria induce distinct statistical dependencies on arm distributions, affecting learning dynamics.

Experimental results

Research questions

  • RQ1How can a unified framework be developed for risk-aware multi-armed bandits across diverse risk criteria?
  • RQ2What general conditions on reward distributions and risk metrics yield a tractable oracle rule?
  • RQ3How do different risk criteria—such as CVaR, mean-variance, and Sharpe ratio—interact with statistical properties of arm rewards?
  • RQ4What are the implications of these differences for designing UCB-based learning policies?
  • RQ5In what ways do risk criteria lead to non-intuitive learning behaviors in stochastic bandits?

Key findings

  • The proposed conditions enable a systematic characterization of the oracle rule across multiple risk criteria, providing a benchmark for regret analysis.
  • Different risk criteria induce distinct statistical dependencies on arm reward distributions, leading to non-intuitive differences in learning behavior.
  • The framework supports the design of UCB policies tailored to risk-aware objectives, such as CVaR and Sharpe ratio, through general principles.
  • Mean-variance and Sharpe ratio criteria exhibit different sensitivity profiles to higher-order moments, affecting exploration-exploitation trade-offs.
  • Conditional value-at-risk (CVaR) requires careful handling of tail behavior, which the framework formalizes through distributional assumptions.
  • The analysis reveals that risk criteria are not interchangeable in bandit learning, even when they appear similar in objective form.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.