Skip to main content
QUICK REVIEW

[论文解读] District Cooling System Control for Providing Operating Reserve based on Safe Deep Reinforcement Learning

Peipei Yu, Hongxun Hui|arXiv (Cornell University)|Dec 21, 2021
Smart Grid Energy Management被引用 5
一句话总结

本文提出了一种无模型的安全深度强化学习(safe-DRL)框架,用于区域供冷系统(DCS),以在电力系统中提供运行备用。通过集成一个安全层以强制执行功率约束,并设计一种自适应奖励函数以防止恢复过程中的功率反弹,该方法在确保热舒适性和系统安全的同时,维持了功率上限合规性并实现了平稳恢复,该方法在真实DCS的数值研究中得到验证。

ABSTRACT

Heating, ventilation, and air conditioning (HVAC) systems are well proved to be capable to provide operating reserve for power systems. As a type of large-capacity and energy-efficient HVAC system (up to 100 MW), district cooling system (DCS) is emerging in modern cities and has huge potential to be regulated as a flexible load. However, strategically controlling a DCS to provide flexibility is challenging, because one DCS services multiple buildings with complex thermal dynamics and uncertain cooling demands. Improper control may lead to significant thermal discomfort and even deteriorate the power system's operation security. To address the above issues, we propose a model-free control strategy based on the deep reinforcement learning (DRL) without the requirement of accurate system model and uncertainty distribution. To avoid damaging "trial & error" actions that may violate the system's operation security during the training process, we further propose a safe layer combined to the DRL to guarantee the satisfaction of critical constraints, forming a safe-DRL scheme. Moreover, after providing operating reserve, DCS increases power and tries to recover all the buildings' temperature back to set values, which may probably cause an instantaneous peak-power rebound and bring a secondary impact on power systems. Therefore, we design a self-adaption reward function within the proposed safe-DRL scheme to constrain the peak-power effectively. Numerical studies based on a realistic DCS demonstrate the effectiveness of the proposed methods.

研究动机与目标

  • 为解决在冷却需求不确定的情况下,对大规模、复杂DCS进行实时运行备用提供控制挑战。
  • 通过在DRL训练期间防止不安全的“试错”行为,确保系统运行安全。
  • 通过避免在DCS提供备用服务后的恢复阶段出现峰值功率反弹,减轻对电力系统的次生影响。
  • 在功率降低和恢复阶段均保持异质建筑内的热舒适性。
  • 开发一种无需精确系统模型或对不确定性分布先验知识的无模型控制策略。

提出的方法

  • 训练一个深度强化学习(DRL)智能体,通过调节DCS的冷冻水质量流量,实现功率的实时调节。
  • 在DRL框架中集成一个安全层,以在训练期间强制执行关键约束,如最大允许功率消耗。
  • 设计一种自适应奖励函数,以惩罚恢复阶段功率的快速上升,从而抑制峰值功率反弹。
  • 该方法为无模型方法,避免依赖DCS的精确热力学模型或对需求不确定性分布的了解。
  • 使用真实DCS的历史数据训练DRL智能体,目标是在保持功率上限的前提下最小化温度偏差。
  • 安全-DRL框架确保DCS在整个功率降低和恢复阶段均运行在安全限值内。

实验结果

研究问题

  • RQ1无模型DRL方法能否有效控制DCS以提供运行备用,同时保持热舒适性?
  • RQ2如何在DRL训练期间防止不安全行为,以确保系统运行安全?
  • RQ3所提方法在DCS恢复阶段对功率反弹的抑制程度如何?
  • RQ4服务时长和功率上限水平的变化在多大程度上影响DCS提供运行备用的性能?
  • RQ5安全-DRL智能体能否在极少再训练的情况下泛化到类似DCS配置?

主要发现

  • 在所有测试场景中,DCS的运行功率始终稳定低于60 MW的功率上限,确保了备用提供期间的安全性。
  • 所有建筑的温度偏差均保持在设定值±1°C以内,整个调节期间均维持了热舒适性。
  • 安全-DRL方法有效抑制了功率反弹,与PI控制器相比,水和风的质量流量恢复过程更缓慢、更平稳。
  • 当服务时长超过30分钟时,温度偏差增加至最高1.5°C,表明在长时间调节下性能出现退化。
  • 当功率上限低于45 MW时,由于物理流量限制,DCS无法满足所需上限,温度偏差显著增加。
  • 该方法在各种时长和功率上限场景下表现出鲁棒性,最优性能在中等时长和较高功率上限下观察到。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。