Skip to main content
QUICK REVIEW

[論文レビュー] Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

Yuto Shibata, Kashu Yamazaki|arXiv (Cornell University)|Mar 11, 2026
Social Robot Interaction and HRI被引用数 0
ひとこと要約

要約: 本論文は、多エージェント強化学習フレームワーク(AssistMimic)を形式化し、近接した人間同士の相互作用に対して物理的知識を取り入れた追従ベースの制御を学習させ、アシスタントと被介助者が共有物理シミュレータ内で共適応し、支援的な運動模倣を達成する。

ABSTRACT

Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking within physics engines (GMT) have enabled virtual characters and humanoid robots to reproduce a broad range of human motions, these behaviors are primarily limited to contact-less social interactions or isolated movements. Assistive scenarios, by contrast, require continuous awareness of a human partner and rapid adaptation to their evolving posture and dynamics. In this paper, we formulate the imitation of closely interacting, force-exchanging human-human motion sequences as a multi-agent reinforcement learning problem. We jointly train partner-aware policies for both the supporter (assistant) agent and the recipient agent in a physics simulator to track assistive motion references. To make this problem tractable, we introduce a partner policies initialization scheme that transfers priors from single-human motion-tracking controllers, greatly improving exploration. We further propose dynamic reference retargeting and contact-promoting reward, which adapt the assistant's reference motion to the recipient's real-time pose and encourage physically meaningful support. We show that AssistMimic is the first method capable of successfully tracking assistive interaction motions on established benchmarks, demonstrating the benefits of a multi-agent RL formulation for physically grounded and socially aware humanoid control.

研究の動機と目的

  • reactive 力の交換を必要とする密接に相互作用する人間の介護・支援シナリオの解決動機づけ
  • 物理シミュレータでのパートナー認識ポリシーを同時に訓練するMARLフレームワークの開発
  • 学習探索と効率を改善するための単一人物モーション priors からの初期化の活用
  • ノイズの多い参照下で安定かつ物理的に意味のある支援を維持するための動的リファレンス再ターゲティングと接触促進報酬の導入

提案手法

  • 物理ベースの人間対人間模倣を、二代理戦の非対称ダイナミクスを持つ有限ホライゾンの多エージェントMDPとして定式化(S=Supporter、R=Recipient)
  • パートナー認識入力と補助的支援状態を備えた単一人物追従制御を拡張して共同学習を可能にする
  • 新しい入力に対してゼロパディングを用い、単一人物モーション priors からポリシーを初期化して学習をブートストラップ
  • 受給者が参照から外れた場合に有効な相対的な手のターゲットを保持する動的リファレンス再ターゲティングの実装
  • 近接接触シナリオで厳密な運動追従よりも積極的で力覚を伴う相互作用を優先する接触促進報酬の導入
  • より広いモーションカバーのために専門家ポリシーを訓練し、DAggerを介してジェネラリストへ蒸留する
Figure 2 : Learning contact-rich assistive behaviors is substantially more difficult in the close-contact interactions that we target (bottom) than in contact-less social interactions (top) or isolated motions ( gray SR curve). AssistMimic addresses these challenges, achieving successful imitation f
Figure 2 : Learning contact-rich assistive behaviors is substantially more difficult in the close-contact interactions that we target (bottom) than in contact-less social interactions (top) or isolated motions ( gray SR curve). AssistMimic addresses these challenges, achieving successful imitation f

実験結果

リサーチクエスチョン

  • RQ1 支援者と被介助者の両方のジョイントMARL訓練は、物理的に一貫し力を交換する支援運動を学習できるか?
  • RQ2 ダイナミックリファレンス再ターゲティングは密接接触支援における接触のロバスト性と安定性を向上させるか?
  • RQ3 単一人物モーション priors からの初期化と接触促進報酬の追加が学習効率と模倣忠実度に与える影響は?
  • RQ4 学習したポリシーは未見の受給者ダイナミクスや生成された相互作用軌道にどれだけロバストか?
  • RQ5 専門家ポリシーをDAggerでジェネラリストへ蒸留することは、様々な相互作用クリップへの一般化を改善するか?

主な発見

  • AssistMimic は Inter-X データセットの成功率が高く、安定性も向上(Inter-X SR 83%、HHI-Assist SR 66%)。
  • ジョイントMARL訓練は、密接に相互作用し力を交換する運動に対して、連続的または切り離した学習手法より優れている。
  • モーション priors からの初期化は収束に不可欠で、これがないと学習が失敗するか報酬を利用する。
  • 動的リファレンス再ターゲティングは有効な相互作用ターゲットを維持し、特にHHI-Assistでロバスト性を向上させる。
  • 接触促進報酬は未見の受給者ダイナミクスへのロバスト性と、ベッド上支援時のCOM安定性を顕著に向上させる。
  • 専門家ポリシーをDAggerで蒸留したジェネラリストは、Inter-Xの多様なクリップに対する性能を改善(SR 64.7% 対 39.8%、蒸留なし)。
  • 本フレームワークは生成モデルの相互作用を追跡・再現でき、生成済みおよび未見のモーションへ広い適用性を示す。
Figure 3 : Overview of AssistMimic . We train tracking-based humanoid control policies for both the recipient and the supporter, optimizing them to imitate a paired reference motion sequence. Our architecture builds on the single-agent tracking framework of PHC [ 10 ] , extending it with partner-awa
Figure 3 : Overview of AssistMimic . We train tracking-based humanoid control policies for both the recipient and the supporter, optimizing them to imitate a paired reference motion sequence. Our architecture builds on the single-agent tracking framework of PHC [ 10 ] , extending it with partner-awa

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。