Skip to main content
QUICK REVIEW

[논문 리뷰] Collaborative AI Teaming in Unknown Environments via Active Goal Deduction

Zuyuan Zhang, Hanhan Zhou|arXiv (Cornell University)|2024. 03. 22.
Human-Automation Interaction and Safety인용 수 5
한 줄 요약

STUN은 unknown agents의 잠재 보상을 추론하기 위해 커널 밀도 베이지안 역학습(KD-BIL)을 제안하고, 알려지지 않은 환경에서의 시너지 협업을 위한 목표 조건 정책의 제로샷 적응을 가능하게 한다.

ABSTRACT

With the advancements of artificial intelligence (AI), we're seeing more scenarios that require AI to work closely with other agents, whose goals and strategies might not be known beforehand. However, existing approaches for training collaborative agents often require defined and known reward signals and cannot address the problem of teaming with unknown agents that often have latent objectives/rewards. In response to this challenge, we propose teaming with unknown agents framework, which leverages kernel density Bayesian inverse learning method for active goal deduction and utilizes pre-trained, goal-conditioned policies to enable zero-shot policy adaptation. We prove that unbiased reward estimates in our framework are sufficient for optimal teaming with unknown agents. We further evaluate the framework of redesigned multi-agent particle and StarCraft II micromanagement environments with diverse unknown agents of different behaviors/rewards. Empirical results demonstrate that our framework significantly advances the teaming performance of AI and unknown agents in a wide range of collaborative scenarios.

연구 동기 및 목표

  • 잠재 목표나 보상을 가진 미지의 에이전트와의 시너지 팀링 필요성에 대한 동기 부여.
  • 관찰된 궤적으로부터 잠재 보상을 추정하기 위한 샘플 효율적인 활성 목표 추론 방법 개발.
  • 未知 에이전트와의 인터페이스 시 제로샷 적응을 지원하기 위해 대리 모델에서 목표 조건 정책을 사전 학습한다.
  • 편향되지 않은 보상 추정이 최적 정책 학습에 충분하다는 것을 증명하고, 맞춤형 다중 에이전트 환경에서 개선된 팀 협력 성능을 보여준다.

제안 방법

  • 잠재 보상인 미지의 에이전트와 협력 에이전트(STUN)를 포함하는 dec-POMDP 프레임워크를 도입한다.
  • 커널 밀도 추정치를 사용하여 관찰된 궤적로부터 잠재 보상의 사후분포를 얻기 위해 Kernel Density Bayesian Inverse Learning (KD-BIL)를 제안한다.
  • 보상의 MAP 추정이 최적성에 충분하지 않으며, 편향되지 않은 보상 추정이 Bellman 수렴을 보장한다는 것을 증명한다.
  • 무작위로 샘플링된 보상을 가진 대리 모델을 사용하여 목표 조건 정책 pi(a|o,R)을 사전 학습하여 편향되지 않은 보상 추정을 통한 제로샷 적응을 가능하게 한다.
  • 재학습 없이 거의 최적에 가까운 협업을 달성하기 위해 편향되지 않은 보상 추정에 조건으로 두고 제로샷 적응 규칙 pi(a|o,âR)를 개발한다.
  • 재설계된 MPE/SMAC 환경에서 중앙집중식 사전 학습과 분산 실행으로 확장성을 입증한다.
Figure 1: We consider the problem of enabling synergistic teaming of AI agents with other unknown agents (e.g., human or autonomous agents that could have latent rewards/objectives) in collaborative task environments.
Figure 1: We consider the problem of enabling synergistic teaming of AI agents with other unknown agents (e.g., human or autonomous agents that could have latent rewards/objectives) in collaborative task environments.

실험 결과

연구 질문

  • RQ1실시간으로 잠재 보상을 가진 미지의 에이전트를 효과적으로 추론하고 협력할 수 있는가?
  • RQ2KD-BIL은 제한된 관찰에서 잠재 보상을 추론하는 데 효율적이고 정확한 방법인가?
  • RQ3편향되지 않은 보상 추정을 사용한 제로샷 적응이 최적 또는 근사 최적의 팀 협력 성능을 보장하는가?
  • RQ4다양한 미지의 에이전트와 다양한 환경에서 STUN 에이전트가 baselines와 비교해 어떻게 수행하는가(MPE/SMAC)?

주요 결과

  • KD-BIL은 잠재 보상에 대한 샘플 효율적인 사후 분포를 제공하며 시간에 따라 변화하는 목표에 작동한다.
  • 편향되지 않은 보상 추정은 정책 학습하에 Bellman 수렴과 최적 Q-값에 필요하고 충분하다.
  • 사전 학습된 목표 조건 정책의 제로샷 적응은 다양한 미지 에이전트에 대해 재학습 없이도 근사 최적의 팀 협력을 달성한다.
  • STUN 에이전트는 베이스라인을 능가하고 재설계된 MPE 및 SMAC 과제에서 근사 최적에 가까운 팀 협력을 달성하며, 어려운 맵 및 다양한 잠재 보상을 가진 에이전트를 포함한다.
  • STUN은 변하는 미지의 에이에이전트에 빠르게 적응하며, 일부 경우 도전적인 맵에서 최대 50%의 성능 향상을 보여준다.
Figure 2: An illustration of our proposed framework. STUN agents $\pi_{i}(\cdot|R)$ are pre-trained using surrogate agent models with sampled latent rewards $R$ . To collaborate with unknown agents, they use inverse learning (step b) on the observed trajectories $\{\tau^{u}_{i}\}$ of the unknown age
Figure 2: An illustration of our proposed framework. STUN agents $\pi_{i}(\cdot|R)$ are pre-trained using surrogate agent models with sampled latent rewards $R$ . To collaborate with unknown agents, they use inverse learning (step b) on the observed trajectories $\{\tau^{u}_{i}\}$ of the unknown age

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.