Skip to main content
QUICK REVIEW

[论文解读] MyoDex: A Generalizable Prior for Dexterous Manipulation

Vittorio Caggiano, Sudeep Dasari|arXiv (Cornell University)|Sep 6, 2023
Muscle activation and electromyography studiesEngineering被引用 3
一句话总结

MyoDex 通过在生理上逼真的肌骨骼人体手模型(MyoHand)上利用多任务强化学习,引入了一种可泛化的操作行为先验,实现了对未见过的高接触力任务的快速少样本适应。其任务解决方案数量是蒸馏基线方法的3倍,学习速度更快4倍;并通过 AdroitDex 在24自由度的 AdroitHand 上验证了迁移能力,样本效率比当前最先进方法高出5倍。

ABSTRACT

Human dexterity is a hallmark of motor control. Our hands can rapidly synthesize new behaviors despite the complexity (multi-articular and multi-joints, with 23 joints controlled by more than 40 muscles) of musculoskeletal sensory-motor circuits. In this work, we take inspiration from how human dexterity builds on a diversity of prior experiences, instead of being acquired through a single task. Motivated by this observation, we set out to develop agents that can build upon their previous experience to quickly acquire new (previously unattainable) behaviors. Specifically, our approach leverages multi-task learning to implicitly capture task-agnostic behavioral priors (MyoDex) for human-like dexterity, using a physiologically realistic human hand model - MyoHand. We demonstrate MyoDex's effectiveness in few-shot generalization as well as positive transfer to a large repertoire of unseen dexterous manipulation tasks. Agents leveraging MyoDex can solve approximately 3x more tasks, and 4x faster in comparison to a distillation baseline. While prior work has synthesized single musculoskeletal control behaviors, MyoDex is the first generalizable manipulation prior that catalyzes the learning of dexterous physiological control across a large variety of contact-rich behaviors. We also demonstrate the effectiveness of our paradigms beyond musculoskeletal control towards the acquisition of dexterity in 24 DoF Adroit Hand. Website: https://sites.google.com/view/myodex

研究动机与目标

  • 开发一种可泛化的灵巧操作行为先验,以实现对未见过的、高接触力任务的快速适应。
  • 探究在生理上逼真的手部模型上进行多任务学习,是否能隐式捕捉到类似人类运动控制的、与任务无关的协同模式。
  • 证明所学习的先验可迁移至其他高维机器人系统,而不仅限于人手。
  • 评估所学习的行为先验在通用性与专业化之间的权衡。
  • 在无需运动捕捉数据的前提下,仅依赖强化学习和内在探索,实现灵巧操作。

提出的方法

  • 该方法在57种多样化的高接触力操作任务上,利用 MyoHand 肌骨骼模型进行多任务强化学习,隐式学习一种与任务无关的行为先验 MyoDex。
  • MyoDex 通过共享策略头和任务特定头的深度强化学习进行训练,实现任务间的参数共享与迁移学习。
  • 该方法利用具有23个关节、40多条肌肉及三阶肌肉动力学的生理准确手部模型,以模拟真实的生物力学和接触力。
  • 消融研究对比了在同质与多样化任务分布上训练的先验,以评估任务多样性对泛化能力的影响。
  • 该方法扩展至24自由度的 AdroitHand,通过在与 MyoDex 相同的14项任务上训练对应先验 AdroitDex,展示了跨平台迁移能力。
  • 性能通过在 TCDM 基准中对未见任务进行微调进行评估,测量成功率和样本效率。
Figure 1: Contact rich manipulation behaviors acquired by MyoDex with a physiological MyoHand
Figure 1: Contact rich manipulation behaviors acquired by MyoDex with a physiological MyoHand

实验结果

研究问题

  • RQ1是否可以通过在肌骨骼手部模型上进行多任务强化学习,训练出一个统一策略,作为在多样化、未见任务上实现灵巧操作的可泛化先验?
  • RQ2预训练任务的多样性在多大程度上影响所学习行为先验的泛化能力?
  • RQ3所学习的先验(MyoDex)在新领域、分布外的操作任务上,能在多大程度上实现更快、更高效的样本学习?
  • RQ4该范式是否可成功迁移至其他高维机器人手部系统,如 AdroitHand?
  • RQ5所学习的先验是否表现出类似肌肉协同模式的协调模式,与生物运动控制原理一致?

主要发现

  • 与蒸馏基线相比,使用 MyoDex 的智能体能够解决约3倍多的未见灵巧操作任务。
  • 在少样本泛化设置下,使用 MyoDex 的智能体学习速度比蒸馏基线快4倍。
  • AdroitDex(AdroitHand 的对应先验)在约1000万步内实现了34个未见 TCDM 任务的74.5%成功率,比之前最先进方法(需5000万步)高出5倍的样本效率。
  • 消融研究显示,预训练于多样化任务(MyoDex Alt Diverse)的性能优于预训练于同质任务(MyoDex Alt Homogenous)的模型,后者早期即出现性能饱和。
  • 肌肉协同分析表明,与专家或蒸馏策略相比,MyoDex 学习到更少但更协调的肌肉激活模式,符合生物运动控制的基本原理。
  • 该方法成功在无需运动捕捉数据的前提下,仅通过强化学习和内在探索,学习到真实、高接触力的灵巧操作行为。
Figure 2: MyoHand - Musculoskeletal Hand model (Caggiano et al., 2022a ) . On the left, rendering of the musculoskeletal structure illustrating bone – in gray – and muscle – in red. On the right a skin-like surface for soft contacts is overlaid to the musculoskeletal model.
Figure 2: MyoHand - Musculoskeletal Hand model (Caggiano et al., 2022a ) . On the left, rendering of the musculoskeletal structure illustrating bone – in gray – and muscle – in red. On the right a skin-like surface for soft contacts is overlaid to the musculoskeletal model.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。