[论文解读] Optimal Policies for Observing Time Series and Related Restless Bandit Problems
本文证明,在一类具有高成本观测的线性二次高斯控制问题中,观测时间序列的最优策略为阈值策略:仅当后验方差超过临界阈值时才进行观测。该结果可扩展至非休息性 bandit 问题,确立了明确的 Whittle 指数的存在性,并依赖于一种新颖的验证定理以及与机械词和非线性动力学的联系。
The trade-off between the cost of acquiring and processing data, and uncertainty due to a lack of data is fundamental in machine learning. A basic instance of this trade-off is the problem of deciding when to make noisy and costly observations of a discrete-time Gaussian random walk, so as to minimise the posterior variance plus observation costs. We present the first proof that a simple policy, which observes when the posterior variance exceeds a threshold, is optimal for this problem. The proof generalises to a wide range of cost functions other than the posterior variance. This result implies that optimal policies for linear-quadratic-Gaussian control with costly observations have a threshold structure. It also implies that the restless bandit problem of observing multiple such time series, has a well-defined Whittle index. We discuss computation of that index, give closed-form formulae for it, and compare the performance of the associated index policy with heuristic policies. The proof is based on a new verification theorem that demonstrates threshold structure for Markov decision processes, and on the relation between binary sequences known as mechanical words and the dynamics of discontinuous nonlinear maps, which frequently arise in physics, control and biology.
研究动机与目标
- 建立针对具有高成本、噪声观测的标量高斯时间序列的基于阈值的观测策略的最优性。
- 将该结果扩展至具有高成本观测的线性二次高斯(LQG)控制问题,证明最优策略中存在阈值结构。
- 证明对多个此类时间序列的非休息性 bandit 问题可采用 Whittle 指数求解,且该指数是明确定义的。
- 提供 Whittle 指数的闭式表达式,并评估相关指数策略相对于启发式策略的性能。
提出的方法
- 为马尔可夫决策过程开发了一种新的验证定理,以证明最优策略中存在阈值结构。
- 利用不连续非线性映射的动力学及其与机械词(二进制序列,源于物理学和控制理论)的联系。
- 应用卡尔曼滤波器来建模在不同观测动作下后验均值和方差的转移。
- 通过莫比乌斯变换将观测控制问题简化为关于后验方差的确定性动态规划。
- 通过在特定代价函数下求解动态规划,推导出 Whittle 指数的闭式表达式。
- 采用形式为 $ V(x,v) = Rx^2 + Rv + g(v) $ 的试探解,以在 LQG 控制中解耦控制与观测子问题。
实验结果
研究问题
- RQ1在何种条件下,阈值策略对于最小化高斯时间序列的后验方差与观测成本之和是最优的?
- RQ2在具有高成本观测的 LQG 控制中,最优策略是否在后验方差上表现出阈值结构?
- RQ3对多个此类时间序列的非休息性 bandit 问题是否可通过 Whittle 指数求解,若可,其是否可显式表示为闭式?
- RQ4在仿真中,该指数策略的性能与启发式策略相比如何?
主要发现
- 仅当后验方差超过临界阈值时才进行观测的阈值策略,是最小化后验方差与观测成本之和的最优策略。
- 在具有高成本观测的 LQG 控制中,最优策略为基于后验方差的阈值策略,结合线性反馈控制律。
- 多个时间序列的非休息性 bandit 问题的 Whittle 指数是明确定义的,并可通过所推导的动态规划以闭式计算。
- Whittle 指数被显式推导为系统参数(包括 $ A, B, D, F, \Sigma_Y(a), c(a) $ 和 $ \beta $)的函数。
- 数值比较表明,该指数策略的性能优于启发式策略。
- 证明依赖于一种新颖的验证定理,并在最优控制背景下建立了机械词与不连续非线性映射动力学之间的深刻联系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。