[논문 리뷰] Online Markov Decision Processes with Terminal Law Constraints
본 논문은 알려지지 않은 동역학과 적대적 손실이 있는 온라인 MDP에 대해 리셋 없는 주기적 프레임워크를 제시하고, 주기적 정책을 정의하며, 다중 에이전트 설정에서 서브선형 주기적 후회를 보장하는 알고리즘을 제공합니다.
Traditional reinforcement learning usually assumes either episodic interactions with resets or continuous operation to minimize average or cumulative loss. While episodic settings have many theoretical results, resets are often unrealistic in practice. The infinite-horizon setting avoids this issue but lacks non-asymptotic guarantees in online scenarios with unknown dynamics. In this work, we move towards closing this gap by introducing a reset-free framework called the periodic framework, where the goal is to find periodic policies: policies that not only minimize cumulative loss but also return the agents to their initial state distribution after a fixed number of steps. We formalize the problem of finding optimal periodic policies and identify sufficient conditions under which it is well-defined for tabular Markov decision processes. To evaluate algorithms in this framework, we introduce the periodic regret, a measure that balances cumulative loss with the terminal law constraint. We then propose the first algorithms for computing periodic policies in two multi-agent settings and show they achieve sublinear periodic regret of order $ ilde O(T^{3/4})$. This provides the first non-asymptotic guarantees for reset-free learning in the setting of $M$ homogeneous agents, for $M > 1$.
연구 동기 및 목표
- 알려지지 않은 동역학을 가진 리셋 없는 온라인 MDP에서 주기적 정책을 형식화한다.
- 터미널 분포 제약을 고려한 주기적 후회 지표를 제안한다.
- 다중 에이전트, 적대적 설정에서 주기적 정책을 계산하기 위한 알고리즘을 개발한다.
- M>1 에이전트 시나리오에서 주기적 후회에 대한 비점근 보장을 확립한다.
제안 방법
- 주기적 정책을 rho P_pi = rho를 만족하는 것으로 정의하여 N 단계 후 초기 분포로의 복귀를 보장한다.
- 상태-행동 분포에 대한 일반적인 볼록 손실을 다루기 위해 볼록 RL 프레임워크를 도입한다.
- 알려지지 않은 전이와 적대적 손실을 다루기 위해 보너스 기반 탐색 미러 디센트 알고리즘 (MDPP-K)을 개발한다.
- 보너스를 통해 실현 가능한 부등식으로 변환된 터미널 법칙 제약이 있는 제약된 MDP 구성식을 사용한다.
- rho_t가 알려진 Framework 1과 rho_t를 추정하고 제한된 리셋이 있는 Framework 2로, 둘 다 주기적 후회 경_bound를 산출한다.
- 적절한 조건하에서 tilde-O(T^{3/4}) 차수의 비점근 주기적 후회 경계를 분석하고 증명한다.
실험 결과
연구 질문
- RQ1온라인 MDP에서 알려지지 않은 동역학이 존재할 때 주기적 정책의 존재를 보장하는 조건은 무엇인가?
- RQ2주기적 손실과 터미널 분포 제약 간의 균형을 맞추는 주기적 후회(R_T)를 어떻게 정의하고 최소화할 수 있는가?
- RQ3다중 동질 에이전트의 리셋 없는 학습에 대해 비점근 보장을 갖는 온라인 알고리즘을 설계할 수 있는가?
- RQ4알려지지 않은 전이 동역학이 실행 가능성에 어떤 영향을 미치며, 보너스가 탐색 및 제약 만족을 어떻게 보장하는가?
주요 결과
- 주기적 정책 개념과 rho 방향으로의 에르고딕성을 보장하는 수축 가정 2를 도입한다.
- 누적 손실 차이와 터미널 분포 편차를 결합한 주기적 후회 R_T를 도입한다.
- MDPP-K 알고리즘은 M>1 에이전트에 대해 tilde-O(T^{3/4})의 서브선형 주기적 후회를 달성한다.
- 알려진 rho_t와 추정 rho_t를 가지는 두 프레임워크를 제공하고 각각의 후회 보장을 제시한다.
- 적대적 손실 하에서 제약된 MD 문제의 실행 가능성과 고확률 경계를 보인다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.