Skip to main content
QUICK REVIEW

[논문 리뷰] Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

Alejandro Escontrela, Xue Bin Peng|arXiv (Cornell University)|2022. 03. 28.
Anomaly Detection Techniques and Applications인용 수 5
한 줄 요약

이 논문은 다리 달린 로봇을 위한 강화학습에서 복잡하고 수작업으로 설계된 보상 함수의 대체로, 최소한의 운동 캡처 데이터에서 학습된 악성 운동 사전을 제안한다. 이러한 데이터 기반 스타일 보상으로 정책을 훈련시킴으로써, 자연스럽고 에너지 효율적인 운동을 유도하며 실제 4족 보행 로봇에 효과적으로 전이되는 현실적인 보행 전환을 달성한다. 이는 복잡한 보상 함수나 수작업으로 설계된 대안보다 효율성과 현실성에서 뛰어난 성능을 발휘한다.

ABSTRACT

Training a high-dimensional simulated agent with an under-specified reward function often leads the agent to learn physically infeasible strategies that are ineffective when deployed in the real world. To mitigate these unnatural behaviors, reinforcement learning practitioners often utilize complex reward functions that encourage physically plausible behaviors. However, a tedious labor-intensive tuning process is often required to create hand-designed rewards which might not easily generalize across platforms and tasks. We propose substituting complex reward functions with "style rewards" learned from a dataset of motion capture demonstrations. A learned style reward can be combined with an arbitrary task reward to train policies that perform tasks using naturalistic strategies. These natural strategies can also facilitate transfer to the real world. We build upon Adversarial Motion Priors -- an approach from the computer graphics domain that encodes a style reward from a dataset of reference motions -- to demonstrate that an adversarial approach to training policies can produce behaviors that transfer to a real quadrupedal robot without requiring complex reward functions. We also demonstrate that an effective style reward can be learned from a few seconds of motion capture data gathered from a German Shepherd and leads to energy-efficient locomotion strategies with natural gait transitions.

연구 동기 및 목표

  • 복잡하고 수작업으로 조정된 보상 함수에 의존하지 않고 시뮬레이션에서 현실적이고 전이 가능한 보행 정책를 훈련시키는 데 도전하는 것.
  • 소량의 운동 캡처 데이터에서 학습된 운동 사전가 복잡한 보상 형상화의 효과적인 대체로 작용할 수 있는지 조사하는 것.
  • 악성 운동 사전를 사용해 훈련된 정책의 에너지 효율성과 현실성은 수작업으로 설계된 보상 또는 스타일 정규화 없이 훈련된 정책과 비교하여 평가하는 것.
  • 악성 운동 사전를 사용해 훈련된 정책가 실제 세계 배포에 일반화되어 자연스러운 보행 전환과 안정적인 속도 추적 성능을 유지할 수 있음을 보여주는 것.

제안 방법

  • 소규모의 4.5초 분량의 게르만 셰퍼드 운동 캡처 데이터에서 학습된 스타일 보상으로, GAN 스타일 프레임워크인 악성 운동 사전(AMP)을 활용한다.
  • 특정 작업 보상(예: 속도 추적)과 함께 학습된 AMP 스타일 보상을 조합하여, 작업 성능과 자연스러운 운동을 동시에 유도하는 정책을 훈련시킨다.
  • 생성된 궤적의 분포와 기준 운동 데이터셋 간의 피어슨 발산을 최소화하기 위해 판별자 네트워크를 사용하여, 제시된 운동의 '스타일'을 효과적으로 인코딩한다.
  • 작업 보상 + 악성 스타일 보상의 조합된 보상을 사용해 강화학습을 통해 정책을 훈련시켜, 물리적으로 타당하고 에너지 효율적인 행동을 학습시키도록 한다.
  • 훈련된 정책를 직접 실제 4족 보행 로봇(A1)에 적용하여 실세계 전이성과 성능을 평가한다.
  • 보행 전환, 에너지 소비(Cost of Transport), 속도 추적 정확도를 분석하여 현실성과 효율성을 평가한다.
Figure 1: Training with Adversarial Motion Priors encourages the policy to produce behaviors which capture the essence of the motion capture dataset while satisfying the auxiliary task objective. Only a small amount of motion capture data is required to train the learning system (4.5 seconds in our
Figure 1: Training with Adversarial Motion Priors encourages the policy to produce behaviors which capture the essence of the motion capture dataset while satisfying the auxiliary task objective. Only a small amount of motion capture data is required to train the learning system (4.5 seconds in our

실험 결과

연구 질문

  • RQ1최소한의 운동 캡처 데이터에서 학습된 악성 운동 사전가 복잡하고 수작업으로 설계된 보상 함수를 효과적으로 대체할 수 있는가?
  • RQ2운동 사전를 사용해 훈련된 정책가 표준 또는 복잡한 보상 형상화를 사용한 정책보다 더 자연스러운 보행 전환과 현실적인 보행 패턴을 보여주는가?
  • RQ3운동 사전를 사용해 훈련된 정책가 수작업으로 설계된 또는 비정형 보상보다 더 높은 에너지 효율성(낮은 Cost of Transport)을 달성하는가?
  • RQ4운동 사전를 사용해 훈련된 정책가 실제 세계 배포에 얼마나 잘 일반화되는가? 안정성과 추적 성능을 유지하는가?

주요 결과

  • 악성 운동 사전를 사용해 훈련된 정책는 수작업으로 설계된 보상 함수를 사용한 정책보다 낮은 Cost of Transport(COT)을 달성하여, 더 높은 에너지 효율성을 보였다.
  • 정책는 명령된 속도에 따라 보행 방식을 자연스럽게 전환하는 능력을 습득하였으며, 예를 들어 패싱에서 캐닝터링으로의 전환을 생물학적 4족 보행과 유사하게 수행하였다.
  • 기준 데이터셋에 포함된 4.5초 분량의 운동만으로도, 정책는 데이터셋에 존재하지 않는 속도 범위에 걸쳐도 성공적으로 속도 명령을 추적하였다.
  • A1 로봇에 정책를 실제 세계에서 배포함으로써 안정적이고 자연스러운 보행, 정확한 속도 추적, 에너지 소비의 최소한의 진동을 보여주었다.
  • 필요에 따라 참조 운동에서 벗어나 작업을 완수하기 위해 정책가 의미 있는 변형을 이끌어내는 등, 강건성과 적응성은 입증되었다.
  • AMP와 단순한 작업 보상의 조합은 복잡한 수작업 보상 설계보다 에너지 효율성과 운동의 현실성에서 뛰어난 성능을 보였다.
Figure 2: Key frames, gait pattern, velocity tracking, and energy-efficiency of the robot dog throughout a trajectory A : Key frames of A1 during a canter motion overlaid on a plain background for contrast. B : Gait diagram indicating contact timing and duration for each foot in black. Training with
Figure 2: Key frames, gait pattern, velocity tracking, and energy-efficiency of the robot dog throughout a trajectory A : Key frames of A1 during a canter motion overlaid on a plain background for contrast. B : Gait diagram indicating contact timing and duration for each foot in black. Training with

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.