Skip to main content
QUICK REVIEW

[논문 리뷰] Elastic Decision Transformer

Yueh-Hua Wu, Xiaolong Wang|arXiv (Cornell University)|2023. 07. 05.
Anomaly Detection Techniques and Applications인용 수 5
한 줄 요약

Elastic Decision Transformer (EDT)은 추론 중에 역사 길이를 동적으로 조정하여 궤적 스티칭을 가능하게 함으로써 Decision Transformers를 향상시킨다. 이는 열 劣한 경로에 대해 짧은 역사 길이를, 최적의 경로에 대해선 긴 역사 길이를 선택함으로써 이루어지며, 이는 시퀀스 생성을 향상시키고 D4RL 및 Atari 벤치마크에서 DT와 Q-학습 기반 방법을 모두 능가하는 정책 학습을 가능하게 한다.

ABSTRACT

This paper introduces Elastic Decision Transformer (EDT), a significant advancement over the existing Decision Transformer (DT) and its variants. Although DT purports to generate an optimal trajectory, empirical evidence suggests it struggles with trajectory stitching, a process involving the generation of an optimal or near-optimal trajectory from the best parts of a set of sub-optimal trajectories. The proposed EDT differentiates itself by facilitating trajectory stitching during action inference at test time, achieved by adjusting the history length maintained in DT. Further, the EDT optimizes the trajectory by retaining a longer history when the previous trajectory is optimal and a shorter one when it is sub-optimal, enabling it to "stitch" with a more optimal trajectory. Extensive experimentation demonstrates EDT's ability to bridge the performance gap between DT-based and Q Learning-based approaches. In particular, the EDT outperforms Q Learning-based methods in a multi-task regime on the D4RL locomotion benchmark and Atari games. Videos are available at: https://kristery.github.io/edt/

연구 동기 및 목표

  • 결정 트랜스포머의 궤적 스티칭 한계를 해결하기 위해, 열 劣한 시연로부터 최적의 궤적을 형성할 수 없음을 해결한다.
  • 궤적 품질에 따라 역사 길이를 적응시킴으로써 오프라인 강화학습에서 더 나은 의사결정을 가능하게 한다.
  • 다양한 길이의 역사 입력을 활용함으로써 다중 작업 오프라인 강화학습 환경에서 성능을 향상시킨다.
  • 주요 아키텍처 변경 없이도 Decision Transformers를 향상시킬 수 있는 계산 효율적인 방법을 개발한다.

제안 방법

  • EDT는 기대값 회귀를 사용하여 다양한 역사 길이에 대해 최대 달성 가능한 보상 값을 추정하고, 추론 시 최적의 역사 길이를 식별한다.
  • 모델은 추정된 보상을 최대화하는 역사 길이를 선택함으로써, 열 劣한 궤적을 만날 경우 자신의 맥락을 '새로 고침'할 수 있다.
  • 열 劣한 경로에서는 짧은 역사 길이를 사용하여 나쁜 과거 행동의 영향을 줄이고 탐색 가능성을 높인다.
  • 최적의 경로에서는 긴 역사 길이를 유지함으로써 행동 예측의 안정성과 일관성을 유지한다.
  • 기존의 Decision Transformer 프레임워크와 원활하게 통합되며, 최소한의 계산 오버헤드를 유발한다.
  • 모델이 현재 경로가 열 劣한 경우 더 유리한 미래 궤적으로 전환할 수 있도록 허용함으로써 궤적 스티칭을 가능하게 한다.
Figure 1 : Normalized return with medium-replay datasets. The dotted gray lines indicate normalized return with medium datasets. By achieving trajectory stitching, our method benefits from worse trajectories and learns a better policy.
Figure 1 : Normalized return with medium-replay datasets. The dotted gray lines indicate normalized return with medium datasets. By achieving trajectory stitching, our method benefits from worse trajectories and learns a better policy.

실험 결과

연구 질문

  • RQ1결정 트랜스포머에서 동적 역사 길이가 열 劣한 시연로부터 효과적인 궤적 스티칭을 가능하게 할 수 있는가?
  • RQ2궤적 품질에 따라 역사 길이를 변화시킬 경우 오프라인 강화학습에서 정책 성능에 어떤 영향을 미치는가?
  • RQ3학습 가능한 역사 길이 메커니즘이 다중 작업 오프라인 강화학습에서 고정 길이 역사 접근법을 능가할 수 있는가?
  • RQ4제안된 방법이 Q-학습 기반 및 표준 DT 접근법과 비교해 D4RL 및 Atari와 같은 표준 벤치마크에서 성능을 향상시킬 수 있는가?

주요 결과

  • EDT는 다중 작업 설정에서 D4RL 로케모션 벤치마크에서 Q-학습 기반 방법을 능가하며, 더 뛰어난 샘플 효율성과 정책 품질을 보여준다.
  • 아틀라스 게임에서 EDT는 오프라인 강화학습 방법 중 최고 성능을 기록했으며, DT 및 Q-학습 베이스라인 모두를 능가한다.
  • 이 방법은 궤적 스티칭을 성공적으로 구현하여, 열 劣한 궤적의 더 나은 세그먼트를 조합해 전체적으로 열 劣한 궤적보다 우수한 궤적을 형성할 수 있도록 한다.
  • 가치 함수 추정에 기반한 최적의 역사 길이를 동적으로 선택함으로써 EDT는 다양한 환경에서 성능을 향상시킨다.
  • EDT가 유도하는 계산 오버헤드는 미미하여 기존의 DT 변종과의 통합에 실용적이다.
  • 실험 결과 EDT는 특히 복잡한 다중 작업 오프라인 강화학습 환경에서 표준 Decision Transformers의 성능을 크게 향상시킨다.
Figure 2 : An overview of our Elastic Decision Transformer architecture. $\tilde{R}$ is the prediction of the maximum return. We also show the environments we used in the experiments on the right. We adopt four tasks for D4RL [ 15 ] and $20$ tasks for Atari games.
Figure 2 : An overview of our Elastic Decision Transformer architecture. $\tilde{R}$ is the prediction of the maximum return. We also show the environments we used in the experiments on the right. We adopt four tasks for D4RL [ 15 ] and $20$ tasks for Atari games.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.