Skip to main content
QUICK REVIEW

[Paper Review] Continuous-time multi-armed bandits under random intervention times

Kei Noba, José Luis Pérez|arXiv (Cornell University)|Mar 4, 2026
Advanced Bandit Algorithms Research0 citations
TL;DR

The paper derives explicit Gittins index characterizations for continuous-time multi-armed bandits with random renewal times, including Lévy-driven arms and exponential inter-arrival cases, and proves optimality of the Gittins strategy.

ABSTRACT

This paper examines multi-armed bandits in which actions are taken at random discrete times. The model consists of $J$ independent arms. When an arm is operated, it must remain active for a random duration, modeled by the inter-arrival time of a (possibly arm-dependent) renewal process. For arms evolving as a Lévy process, we provide an explicit characterization of the Gittins index, which is known to yield an optimal strategy. Furthermore, when the inter-arrival times are exponential and the arms evolve as either a spectrally negative Lévy process, a reflected spectrally negative Lévy process, or a diffusion process, the Gittins index is explicitly characterized in terms of the scale function or diffusion characteristics, respectively. Numerical experiments are performed to support the theoretical results.

Motivation & Objective

  • Motivate the study of multi-armed bandits where actions are taken at random times and arms stay active for random renewals.
  • Provide an explicit Gittins index characterization for arms evolving as Lévy processes.
  • Derive explicit Gittins index expressions under exponential renewal times for spectrally negative Lévy, reflected spectrally negative Lévy, and diffusion arms.
  • Show asymptotic and convergence results linking exponential-renewal indices to classical continuous-time indices.

Proposed method

  • Define a multi-armed bandit with J independent arms and arm-specific renewal times.
  • Formulate the discounted reward and the Gittins index as an optimal stopping problem for each arm.
  • Derive a general Gittins index expression for Lévy-driven arms using fluctuation theory.
  • Obtain explicit index formulas under exponential inter-arrival times for spectrally negative Lévy, reflected spectrally negative Lévy, and diffusion processes via scale functions or diffusion characteristics.
  • Prove asymptotic behavior and convergence results, including weak convergence of measures mu^λ to mu^∞ and index convergence.

Experimental results

Research questions

  • RQ1How can the Gittins index be explicitly characterized when arms follow general Lévy processes with random renewal times?
  • RQ2What are the explicit forms of the Gittins index under exponential renewal times for spectrally negative Lévy, reflected spectrally negative Lévy, and diffusion arms?
  • RQ3How does the Gittins index behave as the renewal rate grows large, and does it converge to the classical continuous-time index?
  • RQ4Does the Gittins index policy remain optimal under arm-dependent renewal times?
  • RQ5How can fluctuation theory and scale functions be leveraged to compute the index in these settings?

Key findings

  • The Gittins index strategy is optimal for the continuous-time bandit with random intervention times.
  • An explicit Gittins index characterization is obtained for general Lévy-process arms.
  • With exponential renewal times, closed-form index expressions are derived for spectrally negative, reflected spectrally negative, and diffusion arms in terms of scale functions or diffusion data.
  • The index converges to the continuous-time limit as the exponential rate increases, via a weak convergence result mu^λ ⇒ mu^∞.
  • Asymptotic behavior shows the index tends to the reward function as the renewal rate tends to zero.
  • Numerical experiments are provided to support the theoretical results.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.