[Paper Review] Restless Bandits in Action: Resource Allocation, Competition and Reservation
This paper proposes an asymptotically optimal resource allocation policy for a system of weakly coupled heterogeneous Restless Multi-Armed Bandits (RMABPs) under global capacity constraints. By extending Whittle relaxation and Weber-Weiss asymptotic optimality to a fixed number of active bandits per RMABP—scaled with total bandit count—it proves asymptotic optimality under this novel scaling regime, offering the first such result for this class of constrained RMABP systems.
We study a resource allocation problem with different requests, and with resources of limited capacity shared by multiple requests. We model this problem as a set of heterogeneous Restless Multi-Armed Bandit Problems (RMABPs) connected by the constraints imposed by the resource capacity. Following the idea of Whittle relaxation, and the asymptotic optimality proof of Weber and Weiss, we propose a simple policy for our problem and prove its asymptotic optimality when the RMABPs are weakly coupled with each other. To the best of our knowledge, this is the first work providing asymptotic optimality results for such a resource allocation problem and such a combination of multiple RMABPs. In applying these, we consider the asymptotic optimality in the case of a fixed number of active restless Bandit Processes (BPs) for each RMABP, rather than a fixed ratio of the number of active BPs to the total number of BPs as considered by Weber and Weiss. In particular, we scale the transition rates of the active bandits in proportion to the total number of BPs.
Motivation & Objective
- To address resource allocation in systems with limited-capacity resources shared among multiple heterogeneous requests.
- To model the problem as a set of weakly coupled Restless Multi-Armed Bandits (RMABPs) subject to global capacity constraints.
- To extend the Whittle relaxation framework to this constrained, heterogeneous RMABP setting with a fixed number of active bandits per RMABP.
- To establish asymptotic optimality of the proposed policy under a new scaling regime where transition rates scale with the total number of bandits.
- To provide the first asymptotic optimality result for a system of multiple, heterogeneous RMABPs under capacity constraints with fixed active bandit counts per process.
Proposed method
- Models the resource allocation problem as a set of heterogeneous Restless Multi-Armed Bandits (RMABPs) with shared capacity constraints.
- Applies Whittle relaxation to derive a tractable index-based policy for each RMABP under the capacity constraint.
- Introduces a novel scaling regime where the number of active bandits per RMABP is fixed, while the total number of bandits grows, and scales transition rates proportionally.
- Uses asymptotic analysis to prove that the proposed policy achieves asymptotic optimality under this scaling regime.
- Extends the theoretical framework of Weber and Weiss to a setting where the ratio of active to total bandits is not fixed, but the absolute number of active bandits per RMABP is constant.
- Establishes that the policy's performance converges to the optimal solution as the system size grows under the proposed scaling.
Experimental results
Research questions
- RQ1Can a Whittle relaxation-based policy achieve asymptotic optimality in a system of heterogeneous, weakly coupled RMABPs under global capacity constraints?
- RQ2What scaling regime enables asymptotic optimality when the number of active bandits per RMABP is fixed rather than proportional to the total number of bandits?
- RQ3How does scaling the transition rates of active bandits with the total number of bandits affect the asymptotic performance of the policy?
- RQ4Is the asymptotic optimality result from Weber and Weiss extendable to systems with multiple, heterogeneous RMABPs under capacity constraints and fixed active bandit counts?
- RQ5What is the theoretical performance guarantee of the proposed policy in large-scale systems with fixed active bandit numbers per RMABP?
Key findings
- The proposed policy achieves asymptotic optimality in the limit as the total number of bandits grows, under a novel scaling regime where the number of active bandits per RMABP is fixed.
- The asymptotic optimality is established under a scaling of transition rates proportional to the total number of bandits, which differs from the standard ratio-based scaling of Weber and Weiss.
- This work provides the first asymptotic optimality result for a system of multiple, heterogeneous RMABPs under global capacity constraints with fixed active bandit counts per process.
- The theoretical analysis confirms that the Whittle relaxation approach remains valid and optimal in this new asymptotic regime.
- The results extend the applicability of Whittle's index policy to a broader class of constrained resource allocation problems with fixed active bandit numbers.
- The framework supports practical deployment in large-scale systems where the number of active requests per process is bounded, yet system size grows.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.