[Paper Review] An Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits
This paper proposes an asymptotically optimal index policy for finite-horizon restless multi-armed bandits with multiple pulls per period, using Lagrangian relaxation to decouple the problem into single-arm subproblems. The policy computes indices from optimal solutions of these subproblems and proves asymptotic optimality as the number of arms grows, without requiring indexability, outperforming state-of-the-art heuristics in simulations.
We consider restless multi-armed bandit (RMAB) with a finite horizon and multiple pulls per period. Leveraging the Lagrangian relaxation, we approximate the problem with a collection of single arm problems. We then propose an index-based policy that uses optimal solutions of the single arm problems to index individual arms, and offer a proof that it is asymptotically optimal as the number of arms tends to infinity. We also use simulation to show that this index-based policy performs better than the state-of-art heuristics in various problem settings.
Motivation & Objective
- To develop a computationally tractable policy for finite-horizon restless bandit problems with multiple pulls per period.
- To extend the applicability of index-based policies beyond the infinite-horizon and single-pull settings.
- To remove the need for an indexability condition, which restricts the validity of prior methods like the Whittle index.
- To prove asymptotic optimality in the limit of large numbers of arms and pulls, under a general setting without state-space restrictions.
- To demonstrate superior finite-sample performance compared to existing heuristics through numerical simulations.
Proposed method
- Leverages Lagrangian relaxation to transform the constrained multi-armed bandit problem into a set of decoupled single-arm problems.
- Derives an index for each arm based on the optimal solution of its corresponding single-arm problem under the relaxed constraint.
- Uses fluid limit analysis and convergence arguments to establish asymptotic optimality as the number of arms tends to infinity.
- Employs a state-space aggregation technique to analyze the behavior of the system in the fluid limit, showing convergence of empirical distributions to deterministic flows.
- Applies a dynamic programming formulation to the single-arm problem to compute the index, ensuring computational tractability independent of the number of arms.
- Validates the policy via simulation on benchmark problems, comparing performance against state-of-the-art heuristics.
Experimental results
Research questions
- RQ1Can an index policy be constructed for finite-horizon restless bandits with multiple pulls per period that remains computationally tractable as the number of arms increases?
- RQ2Does such a policy achieve asymptotic optimality in the limit of large numbers of arms and pulls, without requiring indexability?
- RQ3How does the performance of the proposed index policy compare to existing heuristics in finite-sample settings?
- RQ4Can the asymptotic optimality result be established without imposing restrictive conditions on the state space, such as a maximum of three states?
- RQ5Is the proposed index policy robust to general Markovian transition structures and reward functions?
Key findings
- The proposed index policy is asymptotically optimal as the number of arms and pulls per period grow to infinity at the same rate, under general conditions.
- The policy does not require the indexability condition, which is a key limitation of the Whittle index in general settings.
- The asymptotic optimality proof holds regardless of the number of states in each arm’s state space, unlike prior results that required a 3-state restriction.
- Simulations show the policy consistently outperforms state-of-the-art heuristics across diverse problem instances, including those from advertising, clinical trials, and crowdsourced labeling.
- The index computation is based solely on single-arm optimization, making the method computationally scalable to large-scale problems.
- The fluid limit analysis confirms that the empirical distribution of arm states converges almost surely to a deterministic flow, supporting the asymptotic optimality result.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.