Skip to main content
QUICK REVIEW

[论文解读] Energy Management of Multi-mode Plug-in Hybrid Electric Vehicle using Multi-agent Deep Reinforcement Learning

Min Hua, Cetengfei Zhang|arXiv (Cornell University)|Mar 16, 2023
Electric Vehicles and Infrastructure被引用 5
一句话总结

本文提出了一种用于多模式插电式混合动力汽车(PHEVs)全局能量管理的多智能体深度强化学习(MADRL)框架,采用握手策略与相关性比率协调两个DDPG智能体。在软件在环测试中,该方法相较于单智能体DRL实现最高4%的能效节省,相较于基于规则的系统实现23.54%的能效节省,敏感性分析表明学习率是影响最大的因素。

ABSTRACT

The recently emerging multi-mode plug-in hybrid electric vehicle (PHEV) technology is one of the pathways making contributions to decarbonization, and its energy management requires multiple-input and multipleoutput (MIMO) control. At the present, the existing methods usually decouple the MIMO control into singleoutput (MISO) control and can only achieve its local optimal performance. To optimize the multi-mode vehicle globally, this paper studies a MIMO control method for energy management of the multi-mode PHEV based on multi-agent deep reinforcement learning (MADRL). By introducing a relevance ratio, a hand-shaking strategy is proposed to enable two learning agents to work collaboratively under the MADRL framework using the deep deterministic policy gradient (DDPG) algorithm. Unified settings for the DDPG agents are obtained through a sensitivity analysis of the influencing factors to the learning performance. The optimal working mode for the hand-shaking strategy is attained through a parametric study on the relevance ratio. The advantage of the proposed energy management method is demonstrated on a software-in-the-loop testing platform. The result of the study indicates that the learning rate of the DDPG agents is the greatest influencing factor for learning performance. Using the unified DDPG settings and a relevance ratio of 0.2, the proposed MADRL system can save up to 4% energy compared to the single-agent learning system and up to 23.54% energy compared to the conventional rule-based system.

研究动机与目标

  • 为解决多模式PHEVs中解耦的、局部最优的MIMO控制的局限性。
  • 通过多智能体强化学习开发多模式PHEVs的全局最优能量管理策略。
  • 通过握手策略与相关性比率有效协调多个智能体。
  • 通过敏感性分析识别影响学习性能的关键超参数。
  • 在软件在环平台验证所提方法,并展示其卓越的能效表现。

提出的方法

  • 设计了一种基于两个深度确定性策略梯度(DDPG)智能体的多智能体深度强化学习(MADRL)框架,用于管理多模式PHEV中的能量分配。
  • 引入握手策略以实现智能体间的协作,相关性比率控制交互强度。
  • 通过影响学习性能因素的敏感性分析,确立统一的DDPG超参数。
  • 通过参数研究优化相关性比率,以确定最佳协作模式。
  • 在模拟实时车辆动态与能量流的软件在环平台中实现并测试该系统。
  • 该框架实现了多种车辆运行模式下的联合决策,提升了全局能效。

实验结果

研究问题

  • RQ1如何有效将多智能体深度强化学习应用于多模式PHEVs的多输入、多输出(MIMO)能量管理问题?
  • RQ2何种协调机制可实现在复杂MIMO控制系统中多个DDPG智能体的高效协作?
  • RQ3握手策略中的相关性比率如何影响学习性能与能效?
  • RQ4在该框架中,哪一超参数对DDPG智能体的学习性能影响最大?
  • RQ5所提出的MADRL方法在能效节省方面相较于单智能体DRL与基于规则的系统,优势有多大?

主要发现

  • DDPG智能体的学习率被确定为对学习性能影响最大的因素。
  • 采用统一的DDPG设置与0.2的相关性比率,MADRL系统相较于单智能体DRL方法最高实现4%的能效节省。
  • 与传统的基于规则的能量管理方法相比,所提方法最高可降低23.54%的能耗。
  • 采用最优相关性比率的握手策略实现了两智能体之间稳定高效的协调。
  • 软件在环测试验证了MADRL框架在真实PHEV应用中的鲁棒性与可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。