Skip to main content
QUICK REVIEW

[论文解读] Optimal policy evaluation using kernel-based temporal difference methods

Yaqi Duan, Mengdi Wang|arXiv (Cornell University)|Sep 24, 2021
Statistical Methods and Inference参考文献 49被引用 4
一句话总结

本文提出了一种基于核函数的时差学习方法,用于无限时域折扣马尔可夫奖励过程中的最优策略评估,采用正则化核最小二乘时差估计器,其形式简化为求解包含核矩阵的线性系统。关键贡献在于建立了显式依赖于核特征值与贝尔曼残差方差的非渐近 $L^2(\mu)$ 误差界,表明该误差率在最小最大意义下最优,并揭示了有效时域 $H = (1 - \gamma)^{-1}$ 的缩放行为可能并非立方级,具体取决于核函数与问题实例。

ABSTRACT

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized form of the kernel least-squares temporal difference (LSTD) estimate; in the population limit of infinite data, it corresponds to the fixed point of a projected Bellman operator defined by the associated reproducing kernel Hilbert space. The estimator itself is obtained by computing the projected fixed point induced by a regularized version of the empirical operator; due to the underlying kernel structure, this reduces to solving a linear system involving kernel matrices. We analyze the error of this estimate in the $L^2(μ)$-norm, where $μ$ denotes the stationary distribution of the underlying Markov chain. Our analysis imposes no assumptions on the transition operator of the Markov chain, but rather only conditions on the reward function and population-level kernel LSTD solutions. We use empirical process theory techniques to derive a non-asymptotic upper bound on the error with explicit dependence on the eigenvalues of the associated kernel operator, as well as the instance-dependent variance of the Bellman residual error. In addition, we prove minimax lower bounds over sub-classes of MRPs, which shows that our rate is optimal in terms of the sample size $n$ and the effective horizon $H = (1 - γ)^{-1}$. Whereas existing worst-case theory predicts cubic scaling ($H^3$) in the effective horizon, our theory reveals that there is in fact a much wider range of scalings, depending on the kernel, the stationary distribution, and the variance of the Bellman residual error. Notably, it is only parametric and near-parametric problems that can ever achieve the worst-case cubic scaling.

研究动机与目标

  • 为无限时域马尔可夫奖励过程中的基于核函数的策略评估方法提供精确的统计表征。
  • 分析正则化核最小二乘时差估计器在 $L^2(\mu)$-范数下的估计误差,其中 $\mu$ 为平稳分布。
  • 推导显式依赖于核特征值与贝尔曼残差方差的非渐近误差界。
  • 建立最小最大下界,以证明所推导误差率在样本量 $n$ 与有效时域 $H = (1 - \gamma)^{-1}$ 下的最优性。
  • 揭示最坏情况下的立方级缩放($H^3$)仅在参数化或近似参数化设置中出现,而更丰富的核函数可实现显著更优的缩放行为。

提出的方法

  • 该方法采用正则化形式的核最小二乘时差(LSTD)估计,对应于再生核希尔伯特空间(RKHS)中投影贝尔曼算子的不动点。
  • 通过求解涉及核矩阵的线性系统来计算经验估计器,利用表示定理确保计算上的可实现性。
  • 分析基于经验过程理论,以界定经验解与总体解之间在 $L^2(\mu)$ 范数下的估计误差。
  • 误差界在最小假设下推导——仅对奖励函数与总体解提出要求,无需对马尔可夫链转移算子施加结构性假设。
  • 该方法结合核积分算子的特征结构与贝尔曼残差的方差,以细化误差对问题特异性特征的依赖关系。
  • 在马尔可夫奖励过程的子类上推导最小最大下界,以确立所提上界最优性。

实验结果

研究问题

  • RQ1在无限时域 MRP 设置下,正则化核 LSTD 估计器的非渐近 $L^2(\mu)$ 估计误差是多少?
  • RQ2该误差如何依赖于核算子的特征值与贝尔曼残差的方差?
  • RQ3是否可以避免有效时域 $H = (1 - \gamma)^{-1}$ 的最坏情况立方级缩放?若可以,其条件是什么?
  • RQ4所提出的误差界在相关 MRP 问题类上是否为最小最大最优?
  • RQ5在何种条件下,误差缩放会从次立方级过渡到参数化或近似参数化速率?

主要发现

  • 本文建立了 $L^2(\mu)$ 估计误差的非渐近上界,该上界显式依赖于核算子的特征值与贝尔曼残差的方差。
  • 该上界表明,有效时域 $H = (1 - \gamma)^{-1}$ 的缩放并非普遍为立方级;相反,其性能可因核函数与问题实例的不同而显著改善。
  • 所提出的误差界为最小最大最优,经由在 MRP 子类上推导的匹配下界得到验证。
  • 仅在参数化或近似参数化问题中才会出现 $H$ 的立方级缩放($H^3$),表明更丰富的核类可避免这种最坏情况行为。
  • 分析表明,估计误差由核算子的特征值间隙与平稳分布集中程度之间的相互作用所控制,当特征值衰减较慢时可获得更紧的界。
  • 该方法实现了扰动特征向量的常数阶 $\ell_1$-范数,确保在核表示下对扰动具有稳定性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。