Skip to main content
QUICK REVIEW

[Paper Review] When are Kalman-filter restless bandits indexable?

Christopher R. Dance, Tomi Silander|arXiv (Cornell University)|Sep 15, 2015
Advanced Bandit Algorithms Research24 references3 citations
TL;DR

This paper proves that the restless bandit problem associated with a scalar discrete-time Kalman filter is indexable under certain conditions, establishing that the Whittle index is a non-decreasing function of the belief state. The proof leverages Schur-convexity and mechanical words—binary strings linked to palindromes—to demonstrate monotonicity of the index, resolving a long-standing open problem in restless bandit theory for this class of models.

ABSTRACT

We study the restless bandit associated with an extremely simple scalar Kalman filter model in discrete time. Under certain assumptions, we prove that the problem is indexable in the sense that the Whittle index is a non-decreasing function of the relevant belief state. In spite of the long history of this problem, this appears to be the first such proof. We use results about Schur-convexity and mechanical words, which are particular binary strings intimately related to palindromes.

Motivation & Objective

  • To resolve the long-standing open question of whether the restless bandit associated with a scalar discrete-time Kalman filter is indexable.
  • To establish conditions under which the Whittle index is a non-decreasing function of the belief state, ensuring the validity of Whittle's index policy.
  • To provide the first rigorous proof of indexability for this class of continuous-state restless bandits, despite prior extensive study and lack of conclusive results.
  • To apply advanced mathematical tools—mechanical words and Schur-convexity—to a core problem in active learning and sequential decision-making.

Proposed method

  • The authors model the restless bandit problem using a scalar Kalman filter in discrete time, where each arm's state evolves stochastically even when unobserved.
  • They define the belief state as the posterior variance, which serves as the sufficient statistic for the state of each arm.
  • The proof relies on analyzing the derivative of the Whittle index with respect to the belief state, using majorization theory and Schur-convexity to show monotonicity.
  • Mechanical words—specific binary strings with combinatorial properties related to palindromes—are used to characterize the structure of optimal action sequences.
  • The analysis involves constructing matrix expressions for cumulative cost differences and proving non-negativity using weak supermajorization and matrix inequalities.
  • Key inequalities are derived using the Lagrange multiplier relaxation framework, showing that the optimal action policy remains monotone as the cost of activation increases.

Experimental results

Research questions

  • RQ1Under what conditions is the restless bandit problem associated with a scalar Kalman filter indexable?
  • RQ2Is the Whittle index a non-decreasing function of the posterior variance (belief state) in this model?
  • RQ3Can mechanical words and Schur-convexity be used to establish monotonicity of the index in continuous-state restless bandits?
  • RQ4Does the absence of prior proof for this class of problems stem from the complexity of the belief state dynamics and the lack of suitable mathematical tools?
  • RQ5Can the structural properties of mechanical words help characterize optimal policies in active sensing and sequential monitoring problems?

Key findings

  • The paper establishes that the Kalman-filter restless bandit is indexable under the stated assumptions, meaning the Whittle index is a non-decreasing function of the belief state.
  • The proof demonstrates that the derivative of the Whittle index with respect to the belief state is non-negative, confirming monotonicity.
  • Mechanical words are shown to be instrumental in characterizing the structure of optimal action sequences and enabling the majorization arguments.
  • Schur-convexity is applied to prove that cost differences in action sequences remain non-negative, supporting index monotonicity.
  • The analysis confirms that the denominator of the Whittle index expression remains non-negative across all relevant belief states, ensuring well-defined index computation.
  • The result resolves a long-standing open problem in restless bandit theory, providing a theoretical foundation for using Whittle's index policy in Kalman filtering applications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.