[论文解读] Human-like Energy Management Based on Deep Reinforcement Learning and Historical Driving Experiences
本文提出了一种基于深度强化学习(DDPG)的人类驾驶风格能量管理策略,用于串联-并联混合动力电动汽车,其训练数据来源于经验丰富的驾驶员的历史驾驶数据。通过用真实世界驾驶模式替代基于动态规划的专家示范,该方法在保持近似最优性能的同时,实现了更优的燃油经济性与更快的收敛速度。
Development of hybrid electric vehicles depends on an advanced and efficient energy management strategy (EMS). With online and real-time requirements in mind, this article presents a human-like energy management framework for hybrid electric vehicles according to deep reinforcement learning methods and collected historical driving data. The hybrid powertrain studied has a series-parallel topology, and its control-oriented modeling is founded first. Then, the distinctive deep reinforcement learning (DRL) algorithm, named deep deterministic policy gradient (DDPG), is introduced. To enhance the derived power split controls in the DRL framework, the global optimal control trajectories obtained from dynamic programming (DP) are regarded as expert knowledge to train the DDPG model. This operation guarantees the optimality of the proposed control architecture. Moreover, the collected historical driving data based on experienced drivers are employed to replace the DP-based controls, and thus construct the human-like EMSs. Finally, different categories of experiments are executed to estimate the optimality and adaptability of the proposed human-like EMS. Improvements in fuel economy and convergence rate indicate the effectiveness of the constructed control structure.
研究动机与目标
- 开发一种实时、类人化的混合动力电动汽车(HEVs)能量管理策略(EMS),以模仿经验丰富的驾驶员行为。
- 用实际的历史驾驶数据替代传统的基于动态规划(DP)的专家示范,以增强真实感与适应性。
- 在不牺牲最优性的情况下,提升能量管理的燃油经济性与收敛速度。
- 利用真实世界数据,在多种驾驶条件下验证所提出的基于深度强化学习的EMS。
提出的方法
- 构建了面向控制的串联-并联混合动力传动系统模型,作为仿真环境。
- 采用深度确定性策略梯度(DDPG)算法学习最优功率分配策略。
- 使用经验丰富的驾驶员的历史驾驶数据生成专家示范,而非依赖DP求解结果。
- 利用人类驾驶轨迹训练DDPG智能体,以学习类人控制行为。
- 通过历史数据中展现的高质量行为,确保框架的近似最优性。
- 在多种驾驶循环下评估训练好的策略,以评估其适应性与性能。
实验结果
研究问题
- RQ1基于真实人类驾驶数据训练的DRL能量管理系统,是否能在实时应用的前提下实现与动态规划相当的燃油经济性?
- RQ2用真实驾驶员数据替代基于DP的专家示范,对DRL策略的收敛速度与鲁棒性有何影响?
- RQ3类人控制行为在多变驾驶条件下,对能量管理策略的适应性提升程度如何?
- RQ4所提出方法在学习人类驾驶模式的同时,是否能保持近似最优的性能?
主要发现
- 所提出的类人化EMS在燃油经济性方面接近通过动态规划获得的全局最优水平。
- 基于历史驾驶数据训练的DRL智能体收敛速度优于传统DRL方法(无专家示范)。
- 系统在不同驾驶循环中表现出强适应性,表明具备良好的泛化能力。
- 与基于DP的轨迹相比,使用真实驾驶员数据可生成更具真实感与实用性的控制策略。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。