[논문 리뷰] Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
이 논문은 제약된 마르코프 결정 과정(CMDP)를 위한 두 가지 단일 시간 스케일 정책 그래디언트 원-듀얼 알고리즘—정규화된 정책 그래디언트 원-듀얼(RPG-PD) 및 낙관적 정책 그래디언트 원-듀얼(OPG-PD)—을 제안한다. 엔트로피와 이차 정규화를 활용하고 낙관적 그래디언트 업데이트를 적용하여, 정책 반복이 최적의 제약된 정책으로 비점근적, 비선형 수렴성을 확보한다. 이는 CMDP에서 단일 시간 스케일 알고리즘에 대해 이러한 결과를 처음으로 이룬다.
We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-based policy search methods used in practice, the oscillation of policy iterates in these methods has not been fully understood, bringing out issues such as violation of constraints and sensitivity to hyper-parameters. To fill this gap, we employ the Lagrangian method to cast a constrained MDP into a constrained saddle-point problem in which max/min players correspond to primal/dual variables, respectively, and develop two single-time-scale policy-based primal-dual algorithms with non-asymptotic convergence of their policy iterates to an optimal constrained policy. Specifically, we first propose a regularized policy gradient primal-dual (RPG-PD) method that updates the policy using an entropy-regularized policy gradient, and the dual variable via a quadratic-regularized gradient ascent, simultaneously. We prove that the policy primal-dual iterates of RPG-PD converge to a regularized saddle point with a sublinear rate, while the policy iterates converge sublinearly to an optimal constrained policy. We further instantiate RPG-PD in large state or action spaces by including function approximation in policy parametrization, and establish similar sublinear last-iterate policy convergence. Second, we propose an optimistic policy gradient primal-dual (OPG-PD) method that employs the optimistic gradient method to update primal/dual variables, simultaneously. We prove that the policy primal-dual iterates of OPG-PD converge to a saddle point that contains an optimal constrained policy, with a linear rate. To the best of our knowledge, this work appears to be the first non-asymptotic policy last-iterate convergence result for single-time-scale algorithms in constrained MDPs.
연구 동기 및 목표
- 제약된 MDP에 대한 단일 시간 스케일 기반 정책 기반 원-듀얼 방법에서 비점근적, 마지막 반복 수렴 보장이 부족한 문제를 해결하기 위해.
- 이중 시간 스케일 또는 점근적 수렴 가정으로 인해 기존 라그랑주 기반 정책 탐색 방법에서 발생하는 진동 및 제약 위반 문제를 극복하기 위해.
- 단일 시간 스케일 업데이트 체계 하에서 최적의 제약된 정책으로 수렴하는 알고리즘을 개발하기 위해.
- 유한한 근사 오차를 갖는 함수 근사 기법을 사용하여 큰 상태/행동 공간으로의 수렴 보장을 확장하기 위해.
- 동일한 프레임워크 하에서 새로운 낙관적 그래디언트 기반 방법에 대해 선형 수렴성을 확립하기 위해.
제안 방법
- 엔트로피 정규화된 정책 그래디언트를 이용한 원변수 업데이트와 이차 정규화된 그래디언트 상승을 이용한 이중변수 업데이트를 적용한 단일 시간 스케일 알고리즘인 RPG-PD를 제안한다.
- 라그랑주 방법을 통해 제약된 MDP의 정규화된 사다리꼴 최적화 문제를 수식화하여, 문제를 최소최대 최적화 과제로 변환한다.
- 원변수 및 이중변수 변수에 모두 낙관적 그래디언트 업데이트를 적용하여 최적 정책를 포함하는 사다리꼴 점으로의 선형 수렴을 달성하는 OPG-PD를 도입한다.
- 브레그만 발산과 성능 차이 렘마를 사용하여 정책 가치 향상과 수렴 속도를 분석한다.
- 정책 매개변수화에 함수 근사를 통합하고, 근사 오차까지 고려한 수렴성을 증명한다.
- KL 발산과 소프트맥스 매개변수화의 성질을 활용하여 정책 업데이트의 정확한 그래디언트 표현을 유도한다.
![Figure 1: Convergence performance of RPG-PD, OPG-PD, and primal-dual methods. Learning curves of our RPG-PD ( – – ) and OPG-PD ( — ), and NPG-PD [ 23 ] ( – $\cdot$ – ) and PID Lagrangian [ 29 ] ( $\cdot$ $\cdot$ $\cdot$ $\cdot$ ) methods. The horizontal axes mean the policy iterations $\{\pi_{t}\}_{](https://ar5iv.labs.arxiv.org/html/2306.11700/assets/NPG_primal_dual_comparison_reward.png)
실험 결과
연구 질문
- RQ1단일 시간 스케일 정책 기반 원-듀얼 알고리즘이 CMDP에서 최적의 제약된 정책으로 비점근적 마지막 반복 수렴을 달성할 수 있는가?
- RQ2원변수 및 이중변수 업데이트에 정규화를 적용하면 정책 반복이 안정화되고 제약 위반이 방지되는가?
- RQ3단일 시간 스케일 체계에서 낙관적 그래디언트 업데이트가 제약된 MDP에서 선형 수렴 속도를 제공하는가?
- RQ4함수 근사는 대규모 CMDP에서 정책 반복의 수렴에 어떤 영향을 미치는가?
- RQ5정책 매개변수화에 함수 근사를 사용할 경우 수렴 속도와 근사 오차 사이의 상충 관계는 어떠한가?
주요 결과
- RPG-PD는 정책 반복이 최적의 제약된 정책으로 O(1/√T) 속도로 비선형 수렴하며, 이에 해당하는 정규화된 사다리꼴 점으로의 비선형 수렴 속도를 확보한다.
- OPG-PD는 적절한 조건 하에서 최적 정책를 포함하는 사다리꼴 점으로 O(ρ^T) 속도로 선형 수렴을 달성한다. 여기서 ρ ∈ (0,1)이다.
- RPG-PD의 수렴 보장은 함수 근사 설정으로까지 확장되며, 정책 반복은 근사 오차까지 고려하여 비선형 수렴한다.
- 제안된 방법은 제약된 MDP에서 단일 시간 스케일 원-듀얼 알고리즘에 대해 비점근적 마지막 반복 수렴을 보장하는 최초의 방법이다.
- 계산 실험을 통해 제안된 방법이 기준 방법 대비 제약 위반을 줄이고 수렴 안정성을 향상시키는 데 효과적임을 검증하였다.
- 이론적 분석을 통해 OPG-PD의 낙관적 업데이트 메커니즘이 안정적인 반복과 표준 그래디언트 기반 방법보다 빠른 수렴을 보장함을 확인하였다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.