Skip to main content
QUICK REVIEW

[论文解读] Experimental analysis of data-driven control for a building heating system

Giuseppe Tommaso Costanzo, Sandro Iacovella|arXiv (Cornell University)|Jul 13, 2015
Building Energy and Comfort Optimization参考文献 10被引用 4
一句话总结

本文提出一种基于模型辅助的批量强化学习的数据驱动控制方法,具体为使用虚拟支持元组和领域知识塑造的拟合Q-迭代,以在动态电价下优化建筑供暖控制。该方法在仿真中20天内即达到近似最优性能,并在不同室外温度条件下的真实生活实验室中展示了实用且可用的控制策略。

ABSTRACT

Driven by the opportunity to harvest the flexibility related to building climate control for demand response applications, this work presents a data-driven control approach building upon recent advancements in reinforcement learning. More specifically, model assisted batch reinforcement learning is applied to the setting of building climate control subjected to a dynamic pricing. The underlying sequential decision making problem is cast on a markov decision problem, after which the control algorithm is detailed. In this work, fitted Q-iteration is used to construct a policy from a batch of experimental tuples. In those regions of the state space where the experimental sample density is low, virtual support samples are added using an artificial neural network. Finally, the resulting policy is shaped using domain knowledge. The control approach has been evaluated quantitatively using a simulation and qualitatively in a living lab. From the quantitative analysis it has been found that the control approach converges in approximately 20 days to obtain a control policy with a performance within 90% of the mathematical optimum. The experimental analysis confirms that within 10 to 20 days sensible policies are obtained that can be used for different outside temperature regimes.

研究动机与目标

  • 开发一种用于建筑供暖系统的数据驱动控制策略,利用强化学习实现需求响应应用。
  • 通过使用神经网络生成虚拟样本,对实验元组进行扩充,以应对建筑气候控制中数据稀疏的问题。
  • 通过将领域知识整合到学习到的控制策略中,提升策略性能。
  • 通过仿真进行定量评估,并在真实世界的生活实验室环境中进行定性评估。
  • 在动态电价条件下,证明方法可收敛至近似最优性能。

提出的方法

  • 将控制问题建模为马尔可夫决策过程(MDP),以描述建筑气候控制中的序列决策过程。
  • 采用拟合Q-迭代方法,从建筑系统收集的实验数据元组批量中学习控制策略。
  • 在状态空间中数据稀疏的区域,利用人工神经网络生成虚拟支持元组,以提升策略泛化能力。
  • 通过引入领域知识对所得策略进行优化,以增强其鲁棒性和实用性。
  • 通过仿真和真实生活实验室环境中的实际测试验证该方法。
  • 使用动态电价作为奖励信号,以激励降低能源成本。

实验结果

研究问题

  • RQ1能否从有限的实验数据中高效学习到建筑供暖系统中的数据驱动控制策略?
  • RQ2通过神经网络生成的虚拟支持元组,在数据采样稀疏区域对策略性能有何影响?
  • RQ3在多大程度上,领域知识可以提升学习到的强化学习策略的实用性和鲁棒性?
  • RQ4在动态电价条件下,控制策略多快能收敛至近似最优性能?
  • RQ5学习到的策略在真实世界条件下能否在不同室外温度条件下实现泛化?

主要发现

  • 在仿真中,数据驱动控制方法在约20天的训练时间内收敛至数学最优解的90%以内。
  • 在生活实验室中,仅需10至20天即可获得合理且可用的控制策略,适用于多种不同的室外温度条件。
  • 通过神经网络引入的虚拟支持元组,显著提升了状态空间中数据稀疏区域的策略性能。
  • 与原始强化学习输出相比,引入领域知识后得到的策略更具鲁棒性和实用性。
  • 该方法对动态电价信号表现出强适应能力,有效支持了建筑供暖系统的需求响应。
  • 结果证实了将批量强化学习应用于真实世界建筑气候控制的可行性,且具有实际可接受的收敛时间。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。