[论文解读] Convergence and sample complexity of gradient methods for the model-free linear quadratic regulator problem
本文在无模型连续时间线性二次调节器(LQR)问题中,建立了梯度流和梯度下降的指数稳定性与线性收敛性。证明了达到 ϵ-精度所需的仿真时间与总函数评估次数均随 log(1/ϵ) 变化,显著优于先前的界,并展示了两点梯度估计的最优样本复杂度。
Model-free reinforcement learning attempts to find an optimal control action for an unknown dynamical system by directly searching over the parameter space of controllers. The convergence behavior and statistical properties of these approaches are often poorly understood because of the nonconvex nature of the underlying optimization problems and the lack of exact gradient computation. In this paper, we take a step towards demystifying the performance and efficiency of such methods by focusing on the standard infinite-horizon linear quadratic regulator problem for continuous-time systems with unknown state-space parameters. We establish exponential stability for the ordinary differential equation (ODE) that governs the gradient-flow dynamics over the set of stabilizing feedback gains and show that a similar result holds for the gradient descent method that arises from the forward Euler discretization of the corresponding ODE. We also provide theoretical bounds on the convergence rate and sample complexity of the random search method with two-point gradient estimates. We prove that the required simulation time for achieving $\\epsilon$-accuracy in the model-free setup and the total number of function evaluations both scale as $\\log \\, (1/\\epsilon)$.
研究动机与目标
- 分析无模型连续时间LQR问题中基于梯度方法的收敛性与样本复杂度。
- 建立在稳定反馈增益上梯度流动力学的指数稳定性。
- 推导随机搜索中使用两点梯度估计时,仿真时间与函数评估次数的紧致理论界。
- 通过展示 log(1/ϵ) 缩放而非多项式或反幂次依赖,改进先前的样本复杂度结果。
提出的方法
- 通过LQR问题的凸重参数化,分析非凸优化景观。
- 通过常微分方程(ODE)分析梯度流动力学,并证明在稳定反馈增益上的指数稳定性。
- 应用前向欧拉格式,推导出在合适步长下的梯度下降收敛保证。
- 在无真实系统模型访问的前提下,于随机搜索框架中采用两点梯度估计来近似梯度。
- 利用Polyak–Łojasiewicz不等式与李雅普诺夫分析,建立线性收敛速率。
- 推导逆李雅普诺夫算子范数的界,并利用其控制优化问题中类似海森矩阵结构的条件数。
实验结果
研究问题
- RQ1在无模型LQR问题中,稳定反馈增益上的梯度流动力学是否呈现指数收敛?
- RQ2在无模型连续时间LQR设置中,使用固定步长的梯度下降能否实现线性收敛?
- RQ3在无模型LQR问题中,使用两点梯度估计的随机搜索的样本复杂度是多少?
- RQ4在无法精确计算梯度的情况下,所需仿真时间如何随期望精度 ϵ 变化?
- RQ5是否可将总函数评估次数减少至低于 ϵ 的多项式或反幂次依赖?
主要发现
- 在稳定反馈增益上的梯度流动力学表现出指数稳定性,确保全局收敛至最优控制器。
- 使用合适步长的梯度下降线性收敛至最优反馈增益,收敛速率依赖于系统参数及类似海森矩阵结构的条件数。
- 对于采用两点梯度估计的随机搜索方法,达到 ϵ-精度所需的仿真时间与总函数评估次数均呈 O(log(1/ϵ)) 缩放。
- 样本复杂度显著优于先前工作,后者在不同假设下需 O(1/ϵ⁴ log(1/ϵ)) 或 O(1/ϵ²) 次评估。
- 逆李雅普诺夫算子范数在LQR代价的子水平集上一致有界,从而支持全局收敛保证的推导。
- 结果对目标函数值误差与反馈增益矩阵误差均成立,证实了参数空间中鲁棒的收敛性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。