Skip to main content
QUICK REVIEW

[Paper Review] A vector minmax problem for controlled Markov chains

Sameer Kamal|arXiv (Cornell University)|Nov 2, 2010
Advanced Wireless Network Optimization2 references3 citations
TL;DR

This paper proposes a two-time-scale control scheme for finite-state controlled Markov chains under adversarial disturbances, ensuring almost sure convergence of the vector average reward to a desired closed convex set $\bar{D}$ via Blackwell approachability. The scheme uses adaptive, increasing time durations for stationary strategies, leveraging an elementary two-time-scale argument to achieve robust convergence regardless of adversarial actions.

ABSTRACT

The problem of controlling a finite state Markov chain in the presence of an adversary so as to ensure desired performance levels for a vector of objectives is cast in the framework of Blackwell approachability. Relying on an elementary two time scale construction a control scheme is proposed which ensures almost sure convergence to the desired set regardless of the adversarial actions.

Motivation & Objective

  • To address multi-objective control in Markov decision processes under adversarial disturbances, where classical stochastic control theory falls short.
  • To extend Blackwell approachability—previously applied to repeated games—to controlled Markov chains with vector-valued rewards.
  • To develop a control strategy that ensures almost sure convergence of the running average reward to a desired closed convex set $\bar{D}$, irrespective of adversarial actions.
  • To overcome the slow convergence of prior return-time-based schemes by introducing a dynamic, increasing time-scale policy update mechanism.

Proposed method

  • The scheme employs a two-time-scale construction, where the player holds a stationary strategy for a time duration that increases over time.
  • Strategy switches occur at discrete times $s_n$, with the duration between switches governed by a time-scaling function $t(s_n)$ that depends on the current distance to the target set $\bar{D}$.
  • The time scaling is designed such that the duration increases when the average reward is far from $\bar{D}$, and decreases as the system approaches $\bar{D}$, enabling efficient exploration-exploitation trade-off.
  • The convergence proof relies on an elementary two-time-scale result that bounds the rate of change of the average reward vector and ensures it moves toward $\bar{D}$.
  • A key component is the use of the nearest point projection $x_{\bar{D}}$ to define a descent direction, ensuring that the inner product condition in Assumption 1 is satisfied.
  • The analysis uses subsequence arguments and contradiction to show that all limit points of the average reward sequence lie within $\bar{D}$, under almost sure convergence.

Experimental results

Research questions

  • RQ1Can Blackwell approachability be extended to controlled Markov chains with vector-valued rewards and adversarial disturbances?
  • RQ2How can convergence be accelerated in Blackwell approachability for controlled Markov chains compared to return-time-based schemes?
  • RQ3What control policy structure ensures almost sure convergence to a target set $\bar{D}$ under adversarial actions, even when return times are rare?
  • RQ4Can a two-time-scale strategy with increasing time durations achieve robust convergence without relying on regenerative cycles?

Key findings

  • The proposed scheme ensures almost sure convergence of the running average reward vector to the closure of the desired set $\bar{D}$, regardless of the adversary's strategy.
  • Convergence is established under standard ergodicity and convexity assumptions, with the key condition being the existence of a player strategy that ensures a descent direction toward $\bar{D}$ for any point outside it.
  • The time-scaling mechanism ensures that the duration of each strategy is short initially and increases gradually, enabling efficient exploration and exploitation.
  • The proof relies on a two-time-scale argument that controls the rate of change of the average reward and ensures that limit points of the sequence lie within $\bar{D}$.
  • The scheme avoids dependence on rare return times to a fixed state, thus improving convergence speed in large or slowly mixing chains.
  • The result holds even when the adversary acts adversarially, as long as the Blackwell condition (Assumption 1) is satisfied.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.