Skip to main content
QUICK REVIEW

[论文解读] Weak Convergence Properties of Constrained Emphatic Temporal-difference Learning with Constant and Slowly Diminishing Stepsize

Huizhen Yu|arXiv (Cornell University)|Nov 23, 2015
Reinforcement Learning in Robotics参考文献 36被引用 12
一句话总结

本文在离策略强化学习中,建立了带约束的强调时序差分学习(ETD(λ))在常数步长和缓慢衰减步长下的弱收敛性质。在一般离策略条件下,证明了其渐近收敛至最优值函数的邻域,利用弱Feller马尔可夫链理论与随机逼近方法,并给出了固定常数步长下的偏差界。

ABSTRACT

We consider the emphatic temporal-difference (TD) algorithm, ETD($λ$), for learning the value functions of stationary policies in a discounted, finite state and action Markov decision process. The ETD($λ$) algorithm was recently proposed by Sutton, Mahmood, and White to solve a long-standing divergence problem of the standard TD algorithm when it is applied to off-policy training, where data from an exploratory policy are used to evaluate other policies of interest. The almost sure convergence of ETD($λ$) has been proved in our recent work under general off-policy training conditions, but for a narrow range of diminishing stepsize. In this paper we present convergence results for constrained versions of ETD($λ$) with constant stepsize and with diminishing stepsize from a broad range. Our results characterize the asymptotic behavior of the trajectory of iterates produced by those algorithms, and are derived by combining key properties of ETD($λ$) with powerful convergence theorems from the weak convergence methods in stochastic approximation theory. For the case of constant stepsize, in addition to analyzing the behavior of the algorithms in the limit as the stepsize parameter approaches zero, we also analyze their behavior for a fixed stepsize and bound the deviations of their averaged iterates from the desired solution. These results are obtained by exploiting the weak Feller property of the Markov chains associated with the algorithms, and by using ergodic theorems for weak Feller Markov chains, in conjunction with the convergence results we get from the weak convergence methods. Besides ETD($λ$), our analysis also applies to the off-policy TD($λ$) algorithm, when the divergence issue is avoided by setting $λ$ sufficiently large.

研究动机与目标

  • 为解决标准TD(λ)在离策略学习中因行为策略与目标策略不同而长期存在的发散问题。
  • 在常数与缓慢衰减步长调度下,为带约束ETD(λ)建立收敛性保证。
  • 利用随机逼近理论中的弱收敛方法,刻画ETD(λ)迭代的渐近行为。
  • 将这些结果扩展至当λ足够大时的离策略TD(λ),为带约束版本提供新的收敛性质。
  • 分析常数步长设置下的有限样本行为与偏差界,提升实际稳定性与可解释性。

提出的方法

  • 应用随机逼近理论中的弱收敛方法,分析带约束ETD(λ)迭代的极限行为。
  • 利用学习动态所诱导的马尔可夫链的弱Feller性质,建立遍历性与稳定性。
  • 推导出描述平均迭代极限行为的均值ODE(常微分方程)。
  • 引入两种带偏差项的带约束ETD(λ)变体,以在离策略设置中稳定学习并降低方差。
  • 利用弱Feller链的遍历定理,将样本路径行为与投影贝尔曼方程的解联系起来。
  • 通过证明当λ足够大时,收敛性质在离策略TD(λ)中保持等价,将结果扩展至离策略TD(λ)。

实验结果

研究问题

  • RQ1在何种条件下,带常数步长的带约束ETD(λ)会弱收敛至最优值函数的邻域?
  • RQ2与常数步长相比,缓慢衰减步长下ETD(λ)的渐近行为有何差异?
  • RQ3在固定常数步长下,平均迭代与最优解之间的偏差界是多少?
  • RQ4当λ足够大时,ETD(λ)的收敛结果能否扩展至离策略TD(λ)?
  • RQ5底层马尔可夫链的弱Feller性质如何促进带约束ETD(λ)的收敛性分析?

主要发现

  • 带常数步长的带约束ETD(λ)弱收敛至最优解的邻域,其偏差由与步长成正比的项所界定。
  • 对于缓慢衰减步长,该算法在一般离策略条件下弱收敛至最优解。
  • 带约束ETD(λ)平均过程的均值ODE具有唯一的全局吸引平衡点,确保收敛至期望的值函数。
  • 当λ足够大(如λ=1)时,离策略TD(λ)具有相同的收敛结果,将分析范围扩展至ETD(λ)之外。
  • 由学习动态诱导的马尔可夫链的弱Feller性质确保了遍历性,从而可应用遍历定理进行收敛性分析。
  • 该分析为带约束ETD(λ)与离策略TD(λ)提供了新的理论保证,尤其在处理离策略设置中高方差重要性采样方面具有优势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。