[论文解读] Robust Dynamic Locomotion via Reinforcement Learning and Novel Whole Body Controller
该论文通过将强化学习(RL)与一种新型全身运动控制框架(WBLC)相结合,提出了一套鲁棒的仿人机器人动态行走控制框架。RL策略利用相空间规划器(PSP)和倒立摆模型,学习多步行走模式,实现毫秒级实时、抗干扰控制;WBLC通过优先级任务控制与不等式约束二次规划高效计算出力矩指令,在模拟环境中实现了对持续0.1秒、幅值达520 N的冲击力的鲁棒抵抗。
We propose a robust dynamic walking controller consisting of a dynamic locomotion planner, a reinforcement learning process for robustness, and a novel whole-body locomotion controller (WBLC). Previous approaches specify either the position or the timing of steps, however, the proposed locomotion planner simultaneously computes both of these parameters as locomotion outputs. Our locomotion strategy relies on devising a reinforcement learning (RL) approach for robust walking. The learned policy generates multi step walking patterns, and the process is quick enough to be suitable for real-time controls. For learning, we devise an RL strategy that uses a phase space planner (PSP) and a linear inverted pendulum model to make the problem tractable and very fast. Then, the learned policy is used to provide goal-based commands to the WBLC, which calculates the torque commands to be executed in full-humanoid robots. The WBLC combines multiple prioritized tasks and calculates the associated reaction forces based on practical inequality constraints. The novel formulation includes efficient calculation of the time derivatives of various Jacobians. This provides high-fidelity dynamic control of fast motions. More specifically, we compute the time derivative of the Jacobian for various tasks and the Jacobian of the centroidal momentum task by utilizing Lie group operators and operational space dynamics respectively. The integration of RL-PSP and the WBLC provides highly robust, versatile, and practical locomotion including steering while walking and handling push disturbances of up to 520 N during an interval of 0.1 sec. Theoretical and numerical results are tested through a 3D physics-based simulation of the humanoid robot Valkyrie.
研究动机与目标
- 开发一种适用于全尺寸仿人机器人的实时、鲁棒动态行走控制器,能够应对大范围、突发性的干扰。
- 克服以往方法中固定步态周期或步态位置的局限性,通过强化学习联合优化步态周期与脚掌位置。
- 设计一种计算高效的全身控制器(WBLC),整合分层任务优先级与单边接触及摩擦约束。
- 通过高效计算雅可比矩阵的时间导数,实现对快速运动的高保真动态控制,方法基于李群与操作空间动力学。
- 在Valkyrie机器人三维物理仿真环境中,对真实、非预设的干扰条件下验证该框架的有效性。
提出的方法
- 使用相空间规划器(PSP)与线性倒立摆模型,在离线阶段训练强化学习(RL)策略,以缩小搜索空间并加速学习过程。
- RL策略生成基于目标的步态周期与脚掌位置指令,实现实时多步前瞻规划。
- 提出的新型全身运动控制框架(WBLC)结合基于投影的分层控制与维度降低的二次规划(QP),并施加接触不等式约束,包括摩擦力与单边接触力。
- WBLC通过利用李群算子计算雅可比矩阵的高效解析导数,实现对质心动量与操作空间动力学的精确建模,从而计算出力矩指令。
- 通过分层结构维持任务优先级,优先保证平衡、动量与接触力控制,同时遵守物理约束。
- 将学习到的RL策略与WBLC框架集成,实现在全尺寸三维仿人机器人模型上的实时执行,实现鲁棒、动态的行走模式。
实验结果
研究问题
- RQ1强化学习策略是否能够联合优化全仿人机器人在动态行走中的步态周期与脚掌位置,以实现鲁棒性?
- RQ2全身控制器如何在高效处理优先级任务的同时,整合真实的单边接触与摩擦约束?
- RQ3结合基于投影的速度控制与基于QP的约束强制执行的混合控制框架,在动态行走中的性能表现如何?
- RQ4该系统在动态行走过程中,对大范围、瞬时干扰(如520 N持续0.1秒)的抑制能力达到何种程度?
- RQ5将RL与新型WBLC框架集成,是否能在仿真中实现对快速、动态运动的实时、高保真控制?
主要发现
- RL策略成功实现实时学习多步行走模式,可在瞬间规划数百步,远超步态时间尺度。
- WBLC通过结合基于投影的控制与仅依赖接触点数量的低维QP,实现了高计算效率。
- 系统在模拟环境中成功抵御了持续0.1秒、幅值达520 N的干扰,保持了平衡与稳定步态。
- 通过引入基于李群的雅可比矩阵时间导数计算,实现了对快速运动的高保真动态控制,显著提升了精度与响应速度。
- 该框架成功将基于倒立摆的运动规划方法迁移至全仿人机器人模型,同时保持了任务优先级与接触约束的强制执行。
- 所提出的WBLC在不牺牲计算速度的前提下,通过引入摩擦锥与单边接触等不等式约束,显著优于传统基于投影的方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。