[论文解读] An Exponential Lower Bound for Linearly-Realizable MDPs with Constant Suboptimality Gap
本文在具有常数次优性间隙的线性可实现MDP中建立了强化学习的指数级样本复杂度下界,表明即使在有利的间隙假设下,若无额外结构约束,在线强化学习仍难以处理。该结果揭示了在线强化学习与生成模型设置之间存在指数级差异,后者在相同假设下可实现多项式样本复杂度。
A fundamental question in the theory of reinforcement learning is: suppose the optimal $Q$-function lies in the linear span of a given $d$ dimensional feature mapping, is sample-efficient reinforcement learning (RL) possible? The recent and remarkable result of Weisz et al. (2020) resolved this question in the negative, providing an exponential (in $d$) sample size lower bound, which holds even if the agent has access to a generative model of the environment. One may hope that this information theoretic barrier for RL can be circumvented by further supposing an even more favorable assumption: there exists a \emph{constant suboptimality gap} between the optimal $Q$-value of the best action and that of the second-best action (for all states). The hope is that having a large suboptimality gap would permit easier identification of optimal actions themselves, thus making the problem tractable; indeed, provided the agent has access to a generative model, sample-efficient RL is in fact possible with the addition of this more favorable assumption. This work focuses on this question in the standard online reinforcement learning setting, where our main result resolves this question in the negative: our hardness result shows that an exponential sample complexity lower bound still holds even if a constant suboptimality gap is assumed in addition to having a linearly realizable optimal $Q$-function. Perhaps surprisingly, this implies an exponential separation between the online RL setting and the generative model setting. Complementing our negative hardness result, we give two positive results showing that provably sample-efficient RL is possible either under an additional low-variance assumption or under a novel hypercontractivity assumption (both implicitly place stronger conditions on the underlying dynamics model).
研究动机与目标
- 探究在假设存在常数次优性间隙时,线性可实现MDP的指数级样本复杂度下界是否依然成立。
- 确定次优性间隙假设本身是否足以克服Weisz等人(2020)在在线强化学习设置中识别出的困难障碍。
- 探索额外的结构假设(如低方差或超收缩性)是否能恢复样本效率。
- 澄清在线强化学习与生成模型设置之间在满足线性可实现性和次优性间隙条件下的根本性差异。
- 在超越线性可实现性和常数间隙的条件下,建立样本高效强化学习的必要条件。
提出的方法
- 构建了一类具有线性可实现最优Q函数和常数次优性间隙的MDP家族,以证明其困难性。
- 通过从一个困难的bandit问题进行约化,表明任何算法在特征维度d下都需要指数级样本数量。
- 采用一种新颖的超收缩性下最小二乘回归分析方法,在该假设下获得正面结果。
- 改编了Du等人(2019c)和Weisz等人(2020)的技术,将并集界替换为引理6,以控制困难MDP构造中的估计误差。
- 提出两个正面结果:一个基于低方差假设,另一个基于动力学的新型超收缩性条件。
- 证明了用于下界构造的困难MDP家族违反了低方差和超收缩性假设,表明这些假设对效率而言是必要的。
实验结果
研究问题
- RQ1在在线强化学习设置中,假设存在常数次优性间隙时,线性可实现MDP的指数级样本复杂度下界是否仍然成立?
- RQ2仅靠次优性间隙假设是否足以确保在无生成模型访问的情况下,实现在线强化学习的样本效率?
- RQ3在满足线性可实现性和常数间隙的条件下,还需哪些额外的结构假设才能在在线强化学习中实现多项式样本复杂度?
- RQ4在相同假设下,是否存在在线强化学习设置与生成模型设置之间的指数级分离?
- RQ5低方差或超收缩性假设能否被解释为MDP复杂性的自然表征?
主要发现
- 即使在存在常数次优性间隙和线性可实现性的情况下,在线强化学习的指数级样本复杂度下界仍为 $2^{ ilde{ heta}( ext{min} olimits"){d,H})}$。
- 该下界在无生成模型的情况下依然成立,表明在相同假设下,在线设置比生成模型设置指数级更困难。
- 在低方差假设(假设3)下,样本高效强化学习是可能的,但该假设中所需的常数 $C$ 对于困难MDP家族而言是指数级大的。
- 在超收缩性假设(假设4)下,样本高效强化学习同样可能,但该假设中的超收缩性常数 $C_{\text{hyper}}$ 在困难MDP家族中也是指数级大的。
- 用于下界构造的困难MDP家族具有指数级小的最小可达概率,但可被修改为具有完全可达性($\eta_{\text{min}} = 1$),同时仍保持指数级下界。
- 结果表明,低方差和超收缩性假设均不能由线性可实现性和常数次优性间隙假设推出,因此它们对效率而言是必要的。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。