[논문 리뷰] Offline Pre-trained Multi-Agent Decision Transformer: One Big Sequence Model Tackles All SMAC Tasks
이 논문은 다중 에이gent 강화학습(MARL)을 위한 오프라인 프리트레인을 가능하게 하는 트랜스포머 기반의 통합 시퀀스 모델링 프레임워크인 멀티에이전트 디시전 트랜스포머(MADT)를 제안한다. 대규모 오프라인 스타크래프트 II 데이터셋을 활용하여, 다양한 SMAC 작업에서 뛰어난 샘플 효율성과 강력한 피어리스트/제로샷 일반화 성능을 달성하며, 최신 오프라인 RL 방법들을 능가하고 온라인 파인튜닝 시나리오로의 효과적인 전이를 보여준다.
Offline reinforcement learning leverages previously-collected offline datasets to learn optimal policies with no necessity to access the real environment. Such a paradigm is also desirable for multi-agent reinforcement learning (MARL) tasks, given the increased interactions among agents and with the enviroment. Yet, in MARL, the paradigm of offline pre-training with online fine-tuning has not been studied, nor datasets or benchmarks for offline MARL research are available. In this paper, we facilitate the research by providing large-scale datasets, and use them to examine the usage of the Decision Transformer in the context of MARL. We investigate the generalisation of MARL offline pre-training in the following three aspects: 1) between single agents and multiple agents, 2) from offline pretraining to the online fine-tuning, and 3) to that of multiple downstream tasks with few-shot and zero-shot capabilities. We start by introducing the first offline MARL dataset with diverse quality levels based on the StarCraftII environment, and then propose the novel architecture of multi-agent decision transformer (MADT) for effective offline learning. MADT leverages transformer's modelling ability of sequence modelling and integrates it seamlessly with both offline and online MARL tasks. A crucial benefit of MADT is that it learns generalisable policies that can transfer between different types of agents under different task scenarios. On StarCraft II offline dataset, MADT outperforms the state-of-the-art offline RL baselines. When applied to online tasks, the pre-trained MADT significantly improves sample efficiency, and enjoys strong performance both few-short and zero-shot cases. To our best knowledge, this is the first work that studies and demonstrates the effectiveness of offline pre-trained models in terms of sample efficiency and generalisability enhancements in MARL.
연구 동기 및 목표
- 프리트레이닝을 위한 대규모 오프라인 MARL 데이터셋과 벤치마크의 부족을 해결하기 위해.
- 시퀀스 모델링을 통한 오프라인 프리트레이닝이 다중에이전트 RL에서 샘플 효율성과 이식 가능성 향상에 기여하는지 조사하기 위해.
- 다양한 SMAC 작업 간 일반화가 가능한 통합 시퀀스 모델을 개발하기 위해.
- 트랜스포머 기반 모델링이 오프라인 및 온라인 MARL 환경에서 효과적으로 작용하는지 평가하기 위해.
제안 방법
- 자기주의 어텐션 메커니즘을 사용하여 상태, 행동, 보상 시퀀스를 모델링하는 새로운 멀티에이전트 디시전 트랜스포머(MADT) 아키텍처를 제안한다.
- 다양한 스킬 수준을 가진 스타크래프트 II 환경에서의 오프라인 데이터셋을 사용하여 단일 통합 시퀀스 모델을 프리트레이닝한다.
- 프리트레이닝 및 파인튜닝 동안 보상에서의 목표(예: reward-to-go)를 제외하고 상태 임bedding만 입력으로 사용하여 오프라인 및 온라인 데이터 간 분포 이탈을 방지한다.
- 온라인 파인튜닝을 위해 MAPPO를 온라인 정책 최적화 알고리즘으로 사용하여 프리트레이닝된 MADT를 적응시킨다.
- 조건부 시퀀스 모델링을 적용하여 상태 시퀀스를 기반으로 행동을 예측하고, MARL을 조건부 생성 작업으로 간주한다.
- 다섯 가지 SMAC 시나리오에 걸쳐 다중태스크 프리트레이닝을 통해 정책의 일반화 능력을 향상시킨다.
실험 결과
연구 질문
- RQ1단일 오프라인 프리트레이닝된 시퀀스 모델이 피어리스트 및 제로샷 능력을 통해 여러 SMAC 작업에 일반화될 수 있는가?
- RQ2MADT를 사용한 오프라인 프리트레이닝이 온라인 MARL 파인튜닝에서 샘플 효율성을 어떻게 향상시키는가?
- RQ3왜 프리트레이닝된 모델에서 보상에서의 목표(reward-to-go)의 포함이 온라인 파인튜닝 성능에 악영향을 미치는가?
- RQ4프리트레이닝된 트랜스포머의 온라인 적응을 위해 MAPPO를 사용할 경우 다른 정책 그래디언트 방법과 비교해 샘플 효율성에 어떤 영향을 미치는가?
- RQ5왜 오프라인 전용 버전의 MADT는 온라인 파인튜닝 중에 향상되지 않는가?
주요 결과
- MADT는 SMAC 오프라인 벤치마크에서 최신 오프라인 RL 기반 모델(BC, BCQ, CQL, ICQ)을 모두 능가한다.
- 온라인 파인튜닝 시, 프리트레이닝된 MADT는 단지 250만 번의 환경 상호작용으로부터 시작하는 경우보다 높은 수익을 달성한다.
- 예측할 수 없는 지도(예: 3s_vs_4z)에 대해도 프리트레이닝된 MADT 정책이 효과적으로 일반화되며, 강력한 제로샷 전이 성능을 보여준다.
- 입력 특징에서 보상에서의 목표를 제외함으로써 오프라인 및 온라인 데이터 간 분포 이탈이 감소하여 온라인 파인튜닝 성능이 향상된다.
- MAPPO를 사용한 온라인 파인튜닝은 간단한 정책 그래디언트 방법보다 샘플 효율성이 크게 향상된다.
- 오프라인 전용 버전의 MADT는 온라인 파인튜닝 중에 향상되지 않으며, 이는 지도 학습 기반 프리트레이닝에서 보상 획득 동기가 부족하기 때문임을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.