[论文解读] Towards General Function Approximation in Zero-Sum Markov Games
本论文为具有通用函数逼近的双人零和马尔可夫博弈提出了可证明高效的算法,引入了极小极大可区分维数(Minimax Eluder dimension)以衡量函数类的复杂度。其在线性函数设置下将先前工作的遗憾边界改进了√d因子,并通过广义见证秩(generalized witness rank)建立了样本复杂度边界,实现了在解耦与协同设置下使用神经网络和广义线性模型的样本高效学习。
This paper considers two-player zero-sum finite-horizon Markov games with simultaneous moves. The study focuses on the challenging settings where the value function or the model is parameterized by general function classes. Provably efficient algorithms for both decoupled and {coordinated} settings are developed. In the {decoupled} setting where the agent controls a single player and plays against an arbitrary opponent, we propose a new model-free algorithm. The sample complexity is governed by the Minimax Eluder dimension -- a new dimension of the function class in Markov games. As a special case, this method improves the state-of-the-art algorithm by a $\sqrt{d}$ factor in the regret when the reward function and transition kernel are parameterized with $d$-dimensional linear features. In the {coordinated} setting where both players are controlled by the agent, we propose a model-based algorithm and a model-free algorithm. In the model-based algorithm, we prove that sample complexity can be bounded by a generalization of Witness rank to Markov games. The model-free algorithm enjoys a $\sqrt{K}$-regret upper bound where $K$ is the number of episodes.
研究动机与目标
- 弥合具有通用函数逼近的竞争性强化学习中的理论与实践差距。
- 在解耦与协同设置下,为零和马尔可夫博弈开发可证明高效的算法。
- 为竞争环境中函数类引入新的复杂度度量——极小极大可区分维数与广义见证秩。
- 建立与状态空间大小无关的样本复杂度与遗憾边界。
- 将乐观规划原则扩展至具有分布偏移挑战的多智能体马尔可夫博弈。
提出的方法
- 引入极小极大可区分维数作为零和马尔可夫博弈中函数类的新复杂度度量。
- 在解耦设置下提出一种无模型算法,通过交替乐观性与约束集来管理分布偏移。
- 在协同设置下开发一种无模型算法,利用广义见证秩来限制样本复杂度。
- 采用全局乐观性与约束集,结合椭球势论证,以控制估计误差并确保乐观性。
- 应用累积误差控制与椭球势引理,以界定遗憾并确保收敛性。
- 将MDP中的集中与信息增益方法扩展至马尔可夫博弈,以处理多智能体交互。
实验结果
研究问题
- RQ1能否为具有通用函数逼近的零和马尔可夫博弈开发可证明高效的算法?
- RQ2如何在马尔可夫博弈中适应乐观性以避免分布偏移并确保收敛?
- RQ3为表征函数逼近下竞争性强化学习的样本效率,需要哪些新的复杂度度量?
- RQ4极小极大可区分维数与覆盖数和见证秩等现有度量有何关系?
- RQ5所提出的算法能否实现与状态空间大小无关的遗憾与样本复杂度边界?
主要发现
- 在解耦设置下的无模型算法实现了Õ(H√(d_E K log N_F))的遗憾边界,其中d_E为极小极大可区分维数,N_F为函数类的覆盖数。
- 在d维特征的线性函数逼近设置下,遗憾边界相比先前工作改进了√d因子。
- 在协同设置下,无模型算法需Õ(H³W²/ε²)个样本以达到ε-近似纳什均衡,其中W为广义见证秩。
- 在协同设置下的无模型算法的遗憾为Õ(H√(dK log(N_F N_Π))),其中d为极小极大可区分维数的变体,N_Π为策略类的覆盖数。
- 极小极大可区分维数在表格型博弈中为Õ(|S||A|²),在过参数化神经网络中可通过NTK框架下的临界信息增益有界。
- 在广义线性模型下,极小极大可区分维数在标准正则性条件下有界于O(d(c₂/c₁)² log(c₂R/(c₁ε)))。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。