Kyoto University · Engineering
Jun Morimoto 교수의 연구실은 인공지능과 뇌과학의 융합을 기반으로 한 강력한 로봇 제어 기술을 개발하고 있습니다. 특히 강화학습과 딥러닝을 활용한 로버스트 제어 정책 설계, 모델링 오차 및 외부 교란에 대응하는 최소최대 기반 최적화 기법을 핵심으로 하며, 생물학적 보행 전략을 모방한 단순하고 유연한 제어 전략도 개발하고 있습니다. 인간 수준의 지능과 유사한 안정성과 탄력성을 갖춘 로봇 제어 시스템의 실현을 목표로 하고 있습니다.
Figures are computed from collected data and may differ slightly.
Deep learning (DL) and reinforcement learning (RL) methods seem to be a part of indispensable factors to achieve human-level or super-human AI systems. On the other hand, both DL and RL have strong connections with our brain functions and with neuroscientific findings. In this review, we summarize talks and discussions in the "Deep Learning and Reinforcement Learning" session of the symposium, International Symposium on Artificial Intelligence and Brain Science. In this session, we discussed whe
This letter proposes a new reinforcement learning (RL) paradigm that explicitly takes into account input disturbance as well as modeling errors. The use of environmental models in RL is quite popular for both offline learning using simulations and for online action planning. However, the difference between the model and the real environment can lead to unpredictable, and often unwanted, results. Based on the theory of H(infinity) control, we consider a differential game in which a "disturbing" a
We developed a robust control policy design method in high-dimensional state space by using differential dynamic programming with a minimax criterion. As an example, we applied our method to a simulated five link biped robot. The results show lower joint torques from the optimal control policy compared to a hand-tuned PD servo controller. Results also show that the simulated biped robot can successfully walk with unknown disturbances that cause controllers generated by standard differential dyna
Biological systems seem to have a simpler but more robust locomotion strategy than that of the existing biped walking controllers for humanoid robots. We show that a humanoid robot can step and walk using simple sinusoidal desired joint trajectories with their phase adjusted by a coupled oscillator model. We use the center-of-pressure location and velocity to detect the phase of the lateral robot dynamics. This phase information is used to modulate the desired joint trajectories. We do not expli
We show that a humanoid robot can step and walk using simple sinusoidal desired joint trajectories with their phase adjusted by a coupled oscillator model. We use the center of pressure location and velocity to detect the phase of the lateral robot dynamics. This phase information is used to modulate the desired joint trajectories. We applied the proposed control approach to our newly developed human sized humanoid robot and a small size humanoid robot developed by Sony, enabling them to generat
We developed a robust control policy design method in high-dimensional state space by using differential dynamic programming with a minimax criterion. As an example, we applied our method to a simulated five link biped robot. The results show lower joint torques from the optimal control policy compared to a hand-tuned PD servo controller. Results also show that the simulated biped robot can successfully walk with unknown disturbances that cause controllers generated by standard differential dyna
We propose a model-based reinforcement learning (RL) algorithm for biped walking in which the robot learns to appropriately modulate an observed walking pattern. Via-points are detected from the observed walking trajectories using the minimum jerk criterion. The learning algorithm controls the via-points based on a learned model of the Poincare map of the periodic walking pattern. The model maps from a state in the single support phase and the controlled via-points to a state in the next single
We propose a model-based reinforcement learning algorithm for biped walking in which the robot learns to appropriately place the swing leg. This decision is based on a learned model of the Poincare map of the periodic walking pattern. The model maps from a state at the middle of a step and foot placement to a state at next middle of a step. We also modify the desired walking cycle frequency based on online measurements. We present simulation results, and are currently implementing this approach
We propose a learning method for implementing human-like sequential movements in robots. As an example of dynamic sequential movement, we consider the "stand-up" task for a two-joint, three-link robot. In contrast to the case of steady walking or standing, the desired trajectory for such a transient behavior is very difficult to derive. The goal of the task is to find a path that links a lying state to an upright state under the constraints of the system dynamics. The geometry of the robot is su
We propose a model-based reinforcement learning algorithm for biped walking in which the robot learns to appropriately modulate an observed walking pattern. Via-points are detected from the observed walking trajectories using the minimum jerk criterion. The learning algorithm modulates the via-points as control actions to improve walking trajectories. This decision is based on a learned model of the Poincaré map of the periodic walking pattern. The model maps from a state in the single support p
Open papers in the app to read, cite, and organize with AI.