Skip to main content
QUICK REVIEW

[论文解读] Information Theoretic Regret Bounds for Online Nonlinear Control.

Sham M. Kakade, Akshay Krishnamurthy|arXiv (Cornell University)|Oct 1, 2020
Advanced Bandit Algorithms Research被引用 13
一句话总结

本文提出了一种基于下界置信度的连续控制(LC3)算法,用于在未知动力学系统中的在线非线性控制,该系统被建模为再生核希尔伯特空间(RKHS)中的函数。该算法实现了与系统维度无关的近似最优 O(√T) regret 边界,依赖于信息论量,并通过在非线性控制任务中实现有效的探索,显著提升了学习性能。

ABSTRACT

This work studies the problem of sequential control in an unknown, nonlinear dynamical system, where we model the underlying system dynamics as an unknown function in a known Reproducing Kernel Hilbert Space. This framework yields a general setting that permits discrete and continuous control inputs as well as non-smooth, non-differentiable dynamics. Our main result, the Lower Confidence-based Continuous Control (LC3) algorithm, enjoys a near-optimal O(\sqrt{T}) regret bound against the optimal controller in episodic settings, where T is the number of episodes. The bound has no explicit dependence on dimension of the system dynamics, which could be infinite, but instead only depends on information theoretic quantities. We empirically show its application to a number of nonlinear control tasks and demonstrate the benefit of exploration for learning model dynamics.

研究动机与目标

  • 解决未知、非线性动力学系统中的在线控制问题,其中动力学被建模为已知再生核希尔伯特空间(RKHS)中的函数。
  • 开发一种控制算法,即使在非光滑、不可微的动力学下,也能在分段设置中实现近似最优的 regret。
  • 通过依赖信息论量而非显式依赖系统维度,消除 regret 边界对系统维度的依赖。
  • 通过实证验证探索策略在学习复杂非线性系统动力学方面的有效性。

提出的方法

  • 将未知系统动力学建模为已知再生核希尔伯特空间(RKHS)中的未知函数,从而灵活表示非线性和非光滑的动力学。
  • 设计 LC3 算法,利用下界置信度来平衡序列控制决策中的探索与利用。
  • 基于信息论量(如观测与系统动力学之间的互信息)推导 regret 边界。
  • 确保 regret 边界以 O(√T) 的形式增长,且不显式依赖于动力学的维度,即使维度为无穷大亦成立。
  • 集成探索策略,通过主动探测动力学中不确定区域来提升模型学习效果。
  • 将算法应用于多种非线性控制任务,以实证方式证明其有效性。

实验结果

研究问题

  • RQ1我们能否在不显式依赖系统维度的情况下,实现在在线非线性控制中的近似最优 regret?
  • RQ2信息论量如何被用于约束未知、非线性动力学系统中的 regret?
  • RQ3在非光滑、不可微系统中,结构化探索在学习精确动力学方面起到什么作用?
  • RQ4基于核的函数空间模型能否有效表示复杂非线性动力学,同时支持高效控制?

主要发现

  • LC3 算法实现了 O(√T) 的 regret 边界,这是分段在线控制设置下的近似最优结果。
  • 该 regret 边界不显式依赖于系统动力学的维度,即使在无限维 RKHS 下也成立。
  • 该边界基于信息论量(如互信息)推导得出,反映了系统动力学中的不确定性。
  • 实证结果表明,LC3 中的探索策略显著提升了对非线性系统动力学的学习效果。
  • 该框架成功处理了离散和连续控制输入,以及非光滑和不可微的动力学。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。