[논문 리뷰] Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
이 논문은 노이즈 주입된 행동 복제를 통해 보상 함수를 학습하는 데 자동으로 순위를 매기는 D-REX를 제시하여, 추가 감독 없이도 시연자보다 우수한 정책을 가능하게 한다. 시연보다 나은 imitation이 가능할 때에 대한 이론을 제공하고 MuJoCo 및 Atari 벤치마크에서 강력한 실험적 이득을 보여준다.
The performance of imitation learning is typically upper-bounded by the performance of the demonstrator. While recent empirical results demonstrate that ranked demonstrations allow for better-than-demonstrator performance, preferences over demonstrations may be difficult to obtain, and little is known theoretically about when such methods can be expected to successfully extrapolate beyond the performance of the demonstrator. To address these issues, we first contribute a sufficient condition for better-than-demonstrator imitation learning and provide theoretical results showing why preferences over demonstrations can better reduce reward function ambiguity when performing inverse reinforcement learning. Building on this theory, we introduce Disturbance-based Reward Extrapolation (D-REX), a ranking-based imitation learning method that injects noise into a policy learned through behavioral cloning to automatically generate ranked demonstrations. These ranked demonstrations are used to efficiently learn a reward function that can then be optimized using reinforcement learning. We empirically validate our approach on simulated robot and Atari imitation learning benchmarks and show that D-REX outperforms standard imitation learning approaches and can significantly surpass the performance of the demonstrator. D-REX is the first imitation learning approach to achieve significant extrapolation beyond the demonstrator's performance without additional side-information or supervision, such as rewards or human preferences. By generating rankings automatically, we show that preference-based inverse reinforcement learning can be applied in traditional imitation learning settings where only unlabeled demonstrations are available.
연구 동기 및 목표
- 시연자 성능을 능가하는 모방 학습이 가능하다는 이론적 조건을 제공한다.
- IRL에서 보상 함수의 모호성을 시연의 순위가 감소시킨다는 것을 보여준다.
- 사람의 라벨 없이 자동으로 순위를 생성하는 실용적 방법(D-REX)을 개발한다.
- 시뮬레이션 로봇 공학 및 Atari 벤치마크에서 D-REX를 실험적으로 검증한다.
- D-REX를 표준 모방 학습 및 시연자 성능과 비교한다.
제안 방법
- 라벨이 없는 시연으로부터 정책을 학습하기 위해 행동 복제를 사용한다.
- 복제된 정책에 노이즈를 주입하여 다양한 성능 수준의 궤적을 생성한다.
- 노이즈 수준에서 자동으로 궤적 순위를 도출한다(더 많은 노이즈일수록 성능이 더 나쁜 방향).
- 궤적 순위 보상 외삽(T-REX)을 적용하여 자동 순위에서 보상 함수를 학습한다.
- 학습된 보상 함수를 사용하여 강화 학습으로 정책을 최적화한다.
실험 결과
연구 질문
- RQ1어떤 조건에서 모방 학습이 시연자의 성능을 능가할 수 있는가?
- RQ2노이즈 주입을 통해 자동으로 생성된 라벨 없는 순위가 시연자를 넘어선 일반화 보상을 회복하는 데 충분한 신호를 제공하는가?
- RQ3D-REX가 추가 감독 없이 표준 모방 학습 및 시연자보다 더 나은가?
주요 결과
- D-REX는 MuJoCo 및 Atari 과제에서 종종 시연자보다 나은 성능을 달성하여 대부분의 경우 BC와 GAIL을 능가한다.
- 자동 노이즈 perturbations는 정책 성능의 단조로운 저하를 만들어 내며, (epsilon-greedy 노이즈) 신뢰 가능한 순위를 가능하게 한다.
- 자동 순위를 통해 학습된 보상은 실제 반환과 상관 관계가 있으며 의미 있는 특성을 드러낸다.
- D-REX는 큰 일반화: Atari 과제에서 루프 구멍을 제외하면 평균 약 39%의 개선, MuJoCo 벤치마크에서 상당한 이득.
- 실험에서 D-REX의 최악의 성능은 시연자 또는 표준 IRL 기반 모방 방법보다 낫다.
- D-REX는 보상이나 인간 선호 없이 시연자를 넘어서 일반화하는 최초의 접근법이다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.