[论文解读] Learning Directed Locomotion in Modular Robots with Evolvable Morphologies
本文提出了一种使用贝叶斯优化和HyperNEAT方法,在可演化形态的模块化机器人中学习定向运动的方法。结果表明,与HyperNEAT相比,贝叶斯优化在学习多样化机器人形态的有效控制器方面更具效率,并通过真实世界验证展示了其在存在可测量现实差距的情况下仍能实现足够的轨迹跟踪能力。
We generalize the well-studied problem of gait learning in modular robots in two dimensions. Firstly, we address locomotion in a given target direction that goes beyond learning a typical undirected gait. Secondly, rather than studying one fixed robot morphology we consider a test suite of different modular robots. This study is based on our interest in evolutionary robot systems where both morphologies and controllers evolve. In such a system, newborn robots have to learn to control their own body that is a random combination of the bodies of the parents. We apply and compare two learning algorithms, Bayesian optimization and HyperNEAT. The results of the experiments in simulation show that both methods successfully learn good controllers, but Bayesian optimization is more effective and efficient. We validate the best learned controllers by constructing three robots from the test suite in the real world and observe their fitness and actual trajectories. The obtained results indicate a reality gap that depends on the controllers and the shape of the robots, but overall the trajectories are adequate and follow the target directions successfully.
研究动机与目标
- 解决模块化机器人在非定向步态之外的定向运动问题。
- 研究在多样化模块化机器人形态而非固定结构上控制器的学习方法。
- 评估贝叶斯优化与HyperNEAT在演化随机组合机器人形态控制器方面的有效性。
- 在真实机器人上验证所学习的控制器,并评估现实差距。
提出的方法
- 该研究使用仿真环境,对从预定义模块集合中随机组合形态的模块化机器人进行控制器训练。
- 应用两种学习算法——贝叶斯优化与HyperNEAT,以演化产生神经网络控制器,实现定向运动。
- 适应度函数评估机器人在目标方向上的运动表现,对偏离和低效运动施加惩罚。
- 控制器通过基于前向进展和方向准确性奖励信号的强化学习进行训练。
- 仿真完成后,将表现最佳的控制器部署在三台真实世界的模块化机器人上,以测试其泛化能力。
- 通过比较不同形态和控制器在仿真与物理环境中的轨迹,分析现实差距。
实验结果
研究问题
- RQ1贝叶斯优化能否在多样化模块化机器人形态上有效学习定向运动控制器?
- RQ2贝叶斯优化在学习可演化形态控制器方面与HyperNEAT相比表现如何?
- RQ3在仿真中学习的控制器在多大程度上能泛化到真实世界的模块化机器人?
- RQ4机器人形态在运动性能的现实差距中起到何种影响?
- RQ5控制器设计对物理部署中轨迹精度有何影响?
主要发现
- 在测试的模块化机器人形态系列中,贝叶斯优化在样本效率和最终性能方面均优于HyperNEAT。
- 通过贝叶斯优化学习的控制器在仿真中获得了更高的适应度分数,表明其具备更好的定向运动能力。
- 真实世界实验证实,最佳控制器成功引导机器人沿目标方向运动,尽管与仿真存在可测量的偏差。
- 现实差距因机器人形态和控制器类型而异,部分配置的差异显著大于其他配置。
- 尽管存在现实差距,物理轨迹仍保持充分且方向准确,展示了良好的泛化能力。
- 本研究证实,可演化形态与高效学习算法的结合,可实现模块化机器人中稳健的运动学习。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。