[논문 리뷰] Open-Ended Learning Leads to Generally Capable Agents
이 논문은 XLand를 제시한다. 이는 대규모 다중 작업 3D 환경과 개방형 학습 루프를 통해 제로샷 일반화와 다양한 과제 공간에 걸친 넓은 능력을 갖춘 에이전트를 달성하며 고정 분포 RL을 능가한다. 또한 동적 과제 생성과 정책 증류를 통한 반복적 생성을 통해 지속적 학습과 출현 행동을 촉진한다.
In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We define a universe of tasks within an environment domain and demonstrate the ability to train agents that are generally capable across this vast space and beyond. The environment is natively multi-agent, spanning the continuum of competitive, cooperative, and independent games, which are situated within procedurally generated physical 3D worlds. The resulting space is exceptionally diverse in terms of the challenges posed to agents, and as such, even measuring the learning progress of an agent is an open research problem. We propose an iterative notion of improvement between successive generations of agents, rather than seeking to maximise a singular objective, allowing us to quantify progress despite tasks being incomparable in terms of achievable rewards. We show that through constructing an open-ended learning process, which dynamically changes the training task distributions and training objectives such that the agent never stops learning, we achieve consistent learning of new behaviours. The resulting agent is able to score reward in every one of our humanly solvable evaluation levels, with behaviour generalising to many held-out points in the universe of tasks. Examples of this zero-shot generalisation include good performance on Hide and Seek, Capture the Flag, and Tag. Through analysis and hand-authored probe tasks we characterise the behaviour of our agent, and find interesting emergent heuristic behaviours such as trial-and-error experimentation, simple tool use, option switching, and cooperation. Finally, we demonstrate that the general capabilities of this agent could unlock larger scale transfer of behaviour through cheap finetuning.
연구 동기 및 목표
- 거대한 절차적으로 생성된 환경 내에서 단일 과제 너머로 일반화하는 에이전트의 창출에 대한 동기 부여.
- 월드, 게임, 공동 플레이어 정책을 결합하여 거대하고 매끄럽게 변화하는 과제 공간을 형성하는 환경 공간(XLand)을 정의하고 연구한다.
- 성능의 분위수(percentiles)로 학습을 지속시키기 위해 과제 분포와 목표를 지속적으로 변화시키는 개방형 학습 프로세스를 개발한다.
- 정규화된 점수 분위수로 진전을 측정하고 보류된 평가 과제 전반에서 나타나는 일반적 행동을 분석한다.
제안 방법
- XLand를 소개한다: 제어 가능한 에이전트, 물체, 기믹 및 보상을 갖춘 본질적으로 다중 에이전트, 절차적으로 생성된 3D 세계 공간.
- 과제를 월드, 게임, 공동 플레이어 정책으로 표현하여 다양성과 매끄러운 변화를 갖는 방대한 과제 공간을 형성한다.
- dynamically 생성된 학습 과제를 사용하여 목표를 암시적으로 모델링하는 어텐션 기반 네트워크로 심층 강화학습을 수행한다.
- 세대 간 정책 증류를 통해 새로운 정책을 부트스트랩하고 성능 경계를 재정의하는 개체군 기반의 반복적 학습 체제를 사용한다.
- 평가 공간 전반에서의 정규화된 점수 분위수로 진전을 측정하고 동적 과제 생성과 균일 샘플링을 비교한다.
실험 결과
연구 질문
- RQ1개방형이고 동적으로 생성된 과제 공간에서 학습된 에가전트가 보류된 평가 과제에 대해 제로샷 일반화를 달성할 수 있는가?
- RQ2동적으로 변화하는 학습 과제가 서로 다른 과제의 거대한 공간에서 지속적 학습을 가능하게 하며 고정 분포보다 우수한가?
- RQ3개방형 학습으로 학습된 일반적으로 능력 있는 에이전트에서 어떤 종류의 출현 휴리스틱과 다중 에이전트 행동이 나타나는가?
- RQ4제로샷 학습 이후 이 개방형 프레임워크에서 미세조정이 성능 향상을 어느 정도까지 가능하게 하는가?
주요 결과
- 에이전트는 Hide and Seek, Capture the Flag, Tag 등 다양한 평가 수준에서 제로샷 일반화를 보인다.
- 새로운 과제에 대해 약 1억 단계의 미세조정은 제로샷이나 처음부터 학습하는 것에 비해 현저한 성능 향상을 낳을 수 있다.
- 지시된 탐색, 다른 플레이어를 통한 정보 수집, 협력적 역학과 같은 출현 행동이 평가 시나리오에서 나타난다.
- 과제의 동적 생성은 학습에 결정적이며 과제 공간에서의 균일 샘플링보다 우월하다.
- 에이전트의 일반적 능력은 저렴한 미세조정을 통해 행동의 더 확장 가능한 전이가 가능하다는 잠재력을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.