Skip to main content
QUICK REVIEW

[論文レビュー] Optimal policy evaluation using kernel-based temporal difference methods

Yaqi Duan, Mengdi Wang|arXiv (Cornell University)|Sep 24, 2021
Statistical Methods and Inference参考文献 49被引用数 4
ひとこと要約

本稿は、無限時間割引マルコフ報酬過程における最適方策評価のためのカーネルベースの時系列差分法を提案する。正則化されたカーネル最小二乗時系列差分推定器を用い、カーネル行列を含む線形システムを解くことで、簡略化される。主な貢献は、カーネル固有値とベルマン残差の分散に明示的な依存関係を持つ非漸近的 $L^2(\mu)$ 誤差バインディングを示したことである。このバインディングにより、誤差率がミニマックス的に最適であり、有効ホライズン $H = (1 - \gamma)^{-1}$ におけるスケーリングが、カーネルと問題インスタンスに応じて立方より良好になる可能性があることが明らかになった。

ABSTRACT

We study methods based on reproducing kernel Hilbert spaces for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We study a regularized form of the kernel least-squares temporal difference (LSTD) estimate; in the population limit of infinite data, it corresponds to the fixed point of a projected Bellman operator defined by the associated reproducing kernel Hilbert space. The estimator itself is obtained by computing the projected fixed point induced by a regularized version of the empirical operator; due to the underlying kernel structure, this reduces to solving a linear system involving kernel matrices. We analyze the error of this estimate in the $L^2(μ)$-norm, where $μ$ denotes the stationary distribution of the underlying Markov chain. Our analysis imposes no assumptions on the transition operator of the Markov chain, but rather only conditions on the reward function and population-level kernel LSTD solutions. We use empirical process theory techniques to derive a non-asymptotic upper bound on the error with explicit dependence on the eigenvalues of the associated kernel operator, as well as the instance-dependent variance of the Bellman residual error. In addition, we prove minimax lower bounds over sub-classes of MRPs, which shows that our rate is optimal in terms of the sample size $n$ and the effective horizon $H = (1 - γ)^{-1}$. Whereas existing worst-case theory predicts cubic scaling ($H^3$) in the effective horizon, our theory reveals that there is in fact a much wider range of scalings, depending on the kernel, the stationary distribution, and the variance of the Bellman residual error. Notably, it is only parametric and near-parametric problems that can ever achieve the worst-case cubic scaling.

研究の動機と目的

  • 無限時間割引マルコフ報酬過程におけるカーネルベース方策評価手法の明確な統計的特徴付けを提供すること。
  • 静的分布 $\mu$ における $L^2(\mu)$-ノルムでの正則化されたカーネル最小二乗時系列差分推定器の推定誤差を分析すること。
  • カーネル固有値とベルマン残差の分散に明示的な依存関係を持つ非漸近的誤差バインディングを導出すること。
  • サンプルサイズ $n$ と有効ホライズン $H = (1 - \gamma)^{-1}$ に関して、導出された誤差率のミニマックス最適性を確立すること。
  • 最悪ケースにおける $H$ に対する立方スケーリング ($H^3$) が、パラメトリックまたはニアパラメトリックな設定でのみ発生することを明らかにすること。

提案手法

  • 本手法は、再生カーネルヒルバート空間(RKHS)における射影ベルマン作用素の不動点に対応する、カーネル最小二乗時系列差分(LSTD)推定の正則化形を採用する。
  • 実証的推定器は、カーネル行列を含む線形システムを解くことで計算され、表現定理を活用して計算の実行可能性を保証する。
  • 誤差バインディングの分析には、経験過程理論を用い、経験的および母集団レベルのカーネルLSTD解の間の $L^2(\mu)$ 推定誤差をバインドする。
  • 最小限の仮定の下で分析が行われる——報酬関数と母集団解に関するみたてのみで、マルコフ連鎖の遷移作用素に関する構造的仮定は不要である。
  • カーネル積分作用素の固有構造とベルマン残差の分散を組み込むことで、問題固有の特徴に依存する誤差依存性を精緻化する。
  • 部分クラスのMRP上でのミニマックス下界を導出し、提案された上界の最適性を確立する。

実験結果

リサーチクエスチョン

  • RQ1無限時間割引MRP設定における正則化されたカーネルLSTD推定器の非漸近的 $L^2(\mu)$ 推定誤差は何か?
  • RQ2誤差はカーネル作用素の固有値とベルマン残差の分散にどのように依存するか?
  • RQ3有効ホライズン $H = (1 - \gamma)^{-1}$ における最悪ケースの立方スケーリングを回避できるか?その条件は何か?
  • RQ4提案された誤差バインディングは、関連するMRP問題クラスにおいてミニマックス的に最適か?
  • RQ5誤差スケーリングが、立方より良好なスケーリングからパラメトリックまたはニアパラメトリックレートに移行するのは、どのような条件下か?

主な発見

  • 本稿は、カーネル作用素の固有値とベルマン残差の分散に明示的な依存関係を持つ非漸近的上界を確立した。
  • 上界は、有効ホライズン $H = (1 - \gamma)^{-1}$ におけるスケーリングが、常に立方ではないことを示しており、カーネルと問題インスタンスに応じて顕著に良好になる可能性がある。
  • 提案された誤差バインディングは、部分クラスのMRP上での一致する下界により、ミニマックス的に最適であることが確認された。
  • 立方スケーリング $H^3$ は、パラメトリックまたはニアパラメトリックな問題でのみ発生し、より豊かなカーネルクラスではこの最悪ケース行動を回避できることが示された。
  • 分析により、推定誤差はカーネル作用素の固有ギャップと静的分布の集中度の相互作用によって制御され、固有値の減少が遅い場合にはよりタイトなバインディングが得られることを示した。
  • 本手法は、摂動に対するカーネル表現の安定性を保証する、定数オーダーの摂動固有ベクトルの $\ell_1$-ノルムを達成した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。