[论文解读] The Power of Linear Controllers in LQR Control
本文分析了随机线性二次型调节器(LQR)控制中的策略后悔问题,比较了三种控制器:最优在线(基于Riccati方程)、最优离线线性(状态反馈)以及最优离线(全局最优动作)。研究证明,当时间范围 T → ∞ 时,最优离线线性策略的成本几乎必然收敛至最优在线策略的成本,这意味着相对于最优离线策略存在线性后悔——即使在独立同分布(i.i.d.)噪声下也是如此。这确立了在后悔最小化框架下,线性控制器在LQR控制中的根本性能极限。
The Linear Quadratic Regulator (LQR) framework considers the problem of regulating a linear dynamical system perturbed by environmental noise. We compute the policy regret between three distinct control policies: i) the optimal online policy, whose linear structure is given by the Ricatti equations; ii) the optimal offline linear policy, which is the best linear state feedback policy given the noise sequence; and iii) the optimal offline policy, which selects the globally optimal control actions given the noise sequence. We fully characterize the optimal offline policy and show that it has a recursive form in terms of the optimal online policy and future disturbances. We also show that cost of the optimal offline linear policy converges to the cost of the optimal online policy as the time horizon grows large, and consequently the optimal offline linear policy incurs linear regret relative to the optimal offline policy, even in the optimistic setting where the noise is drawn i.i.d from a known distribution. Although we focus on the setting where the noise is stochastic, our results also imply new lower bounds on the policy regret achievable when the noise is chosen by an adaptive adversary.
研究动机与目标
- 刻画在随机LQR中,最优在线、最优离线线性和最优离线控制策略之间的性能差距。
- 以这三种控制器之间的时间平均成本差异为度量,量化策略后悔。
- 证明尽管受到线性反馈的约束,最优离线线性策略的成本仍会渐近地匹配最优在线策略的成本。
- 通过利用随机设定下的结果,推导在对抗性噪声下策略后悔的新下界。
- 为在线LQR控制中线性控制器在后悔最小化下的局限性提供理论基础。
提出的方法
- 通过Riccati方程推导最优在线策略,得到因果线性反馈控制律 ut = -Ktxt。
- 将最优离线线性策略表征为线性状态反馈控制器 K*,其在已知噪声序列 w 下最小化总成本。
- 利用最优在线策略与未来扰动的关系,对最优离线策略进行递归分解。
- 应用McDiarmid不等式,证明稳定线性策略的时间平均成本几乎必然集中在其无限时域期望值附近。
- 证明当 T → ∞ 时,最优离线线性策略的成本几乎必然收敛至最优在线策略的成本。
- 通过涉及系统矩阵和Riccati方程解的迹表达式,推导出三种策略之间的时间平均策略后悔的显式表达式。
实验结果
研究问题
- RQ1在随机LQR中,最优离线线性策略的成本与最优在线策略的成本在渐近意义上如何比较?
- RQ2当 T → ∞ 时,最优在线策略与最优离线线性策略之间的时间平均策略后悔是多少?
- RQ3能否利用随机设定下的结果,推导出在对抗性噪声下的策略后悔下界?
- RQ4从未来扰动的角度看,最优离线策略与最优在线策略之间存在何种结构性关系?
- RQ5最优离线线性策略相对于最优离线策略是否能实现次线性后悔?
主要发现
- 当 T → ∞ 时,最优离线线性策略的成本几乎必然收敛至最优在线策略的成本,这意味着两者之间的时间平均后悔趋于零。
- 最优在线策略与最优离线策略之间的时间平均策略后悔收敛至一个非零常数,该常数涉及系统矩阵的迹和Riccati方程的解。
- 最优离线线性策略与最优离线策略之间的时间平均策略后悔收敛至与上述相同的非零常数。
- 即使噪声为独立同分布且来自已知分布,最优离线线性策略相对于最优离线策略仍存在线性后悔。
- 结果表明,在对抗性LQR设置下,任何因果策略可实现的策略后悔存在一个根本下界。
- 成本收敛性通过集中不等式以及最优控制器下闭环系统的稳定性得以建立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。