[论文解读] Stochastic Optimal Control of HVAC system for Energy-efficient Buildings
该论文提出了一种基于马尔可夫决策过程(MDPs)和基于梯度的学习(GB-L)方法的随机最优控制框架,用于在天气和人员 occupancy 不确定的情况下,平衡暖通空调(HVAC)系统的能效与热舒适性。该方法在在线计算时间少于1秒、热舒适性概率高的情况下实现了接近最优的性能,尽管与具备完美未来信息的理想模型预测控制相比,能源成本增加了6.5%。
The heating, ventilation and air-conditioning (HVAC) system accounts for substantial energy use in buildings, whereas a large group of occupants are still not actually feeling comfortable staying inside. This poses the issue of developing energy-efficient HVAC control, i.e., reduce energy use (cost) while simultaneously enhancing human comfort. This paper pursues the objective and studies the stochastic optimal HVAC control subject to uncertain thermal demand (i.e., the weather and occupancy etc). Particularly, we involve the elaborate predicted mean vote (PMV) thermal comfort model in the optimization. The problem is computationally challenging due to the non-linear and non-analytical constraints imposed by the system dynamics and PMV model. We make the following contributions to address it. First, we formulate the problem as a Markov decision process (MDP) which is a desirable modeling technique capable of handling the complexities. Second, we propose a gradient-based learning (GB-L) method for progressively learning a stochastic control policy off-line and store it for on-line execution. Third, we prove the learning method converge to the optimal policies theoretically, and its performance (i.e., energy cost, thermal comfort and on-line computation) for HVAC control via simulations. The comparisons with the existing model predictive control based relaxation (MPC-R) method which is assumed with accurate future information and supposed to provide the near-optimal bounds, show that though there exists some performance loss in energy cost reduction (i.e., 6.5%), the proposed method can enable efficient on-line implementation (less than 1 second) and provide high probability of thermal comfort under uncertainties.
研究动机与目标
- 解决在天气和人员 occupancy 引起的不确定热需求下,平衡暖通空调系统能效与居住者热舒适性的挑战。
- 将暖通空调控制建模为马尔可夫决策过程(MDP),以处理非线性动态和复杂约束。
- 开发一种基于梯度的学习(GB-L)方法,用于离线策略学习,实现实时在线执行的快速响应。
- 确保所学习控制策略的理论收敛至全局最优。
- 在不确定性下展示鲁棒性能,具有高热舒适性概率和低在线计算时间。
提出的方法
- 将暖通空调系统建模为具有状态、动作和奖励函数的马尔可夫决策过程(MDP),以捕捉随机动态和热舒适性约束。
- 在MDP框架内集成预测平均投票(PMV)模型作为非解析、非线性的约束,以量化热舒适性。
- 提出一种基于梯度的策略改进(GBPI)算法,通过策略梯度估计迭代更新随机控制策略。
- 采用参数化的随机策略并结合softmax动作选择,以实现探索,并利用梯度进行策略优化。
- 该方法采用基于策略梯度定理推导的性能梯度更新规则,确保收敛至最优策略。
- GB-L方法在离线阶段训练并在在线阶段部署,实现响应时间低于1秒的实时控制。
实验结果
研究问题
- RQ1在不确定的热扰动下,该随机最优控制框架能否有效平衡暖通空调系统的能耗与热舒适性?
- RQ2与假设具备完美未来信息的理想模型预测控制(MPC-R)相比,所提出的基于梯度的学习方法在能耗和舒适性方面表现如何?
- RQ3尽管暖通空调系统和PMV模型存在非凸和非线性约束,GB-L方法是否能收敛至全局最优策略?
- RQ4所提出方法在能耗降低与在线计算效率之间存在何种权衡?
- RQ5该方法能否在真实世界不确定性(如天气和人员 occupancy 变化)下确保高热舒适性概率?
主要发现
- 所提出的GB-L方法与理想MPC-R基线相比,仅增加6.5%的能耗,实现了接近最优的性能。
- 该方法可在每个控制步长内实现少于1秒的在线实现,适用于实时部署。
- 理论分析证明,在给定的MDP建模下,GB-L方法可收敛至全局最优策略。
- 该方法在不确定性下确保了高热舒适性概率,优于传统控制策略在维持居住者满意度方面的表现。
- 仿真结果证实,将PMV模型集成到MDP框架中,能有效平衡能耗与热舒适性。
- 基于梯度的策略更新机制成功应对了暖通空调系统和PMV模型的非凸、非解析约束。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。