Skip to main content
QUICK REVIEW

[論文レビュー] Stochastic Optimal Control of HVAC system for Energy-efficient Buildings

Yu Yang, Guoqiang Hu|arXiv (Cornell University)|Nov 3, 2019
Building Energy and Comfort Optimization参考文献 29被引用数 6
ひとこと要約

本稿では、天候および占有人数の不確実性を考慮したHVACシステムにおけるエネルギー効率性と熱的快適性のバランスを図るために、マーカフ連鎖過程(MDP)と勾配ベース学習(GB-L)手法を用いた確率的最適制御フレームワークを提案する。本手法は、理想的な予測制御(MPC)に比べて6.5%のエネルギーコスト増加があるものの、1秒未塔のオンライン計算時間で近似的最適な性能を達成し、高い熱的快適性確率を実現する。

ABSTRACT

The heating, ventilation and air-conditioning (HVAC) system accounts for substantial energy use in buildings, whereas a large group of occupants are still not actually feeling comfortable staying inside. This poses the issue of developing energy-efficient HVAC control, i.e., reduce energy use (cost) while simultaneously enhancing human comfort. This paper pursues the objective and studies the stochastic optimal HVAC control subject to uncertain thermal demand (i.e., the weather and occupancy etc). Particularly, we involve the elaborate predicted mean vote (PMV) thermal comfort model in the optimization. The problem is computationally challenging due to the non-linear and non-analytical constraints imposed by the system dynamics and PMV model. We make the following contributions to address it. First, we formulate the problem as a Markov decision process (MDP) which is a desirable modeling technique capable of handling the complexities. Second, we propose a gradient-based learning (GB-L) method for progressively learning a stochastic control policy off-line and store it for on-line execution. Third, we prove the learning method converge to the optimal policies theoretically, and its performance (i.e., energy cost, thermal comfort and on-line computation) for HVAC control via simulations. The comparisons with the existing model predictive control based relaxation (MPC-R) method which is assumed with accurate future information and supposed to provide the near-optimal bounds, show that though there exists some performance loss in energy cost reduction (i.e., 6.5%), the proposed method can enable efficient on-line implementation (less than 1 second) and provide high probability of thermal comfort under uncertainties.

研究の動機と目的

  • 天候および占有人数による不確実な熱的需要を考慮したHVACシステムにおけるエネルギー効率性と居住者の熱的快適性のバランスを図る課題に対処する。
  • 非線形ダイナミクスと複雑な制約を扱えるよう、HVAC制御をマーカフ連鎖過程(MDP)として定式化する。
  • オンライン実行を高速化できるように、オフラインでのポリシー学習を可能にする勾配ベース学習(GB-L)手法を開発する。
  • 学習済み制御ポリシーの理論的収束を、グローバル最適解に保証する。
  • 変動する天候や占有人数といった現実世界の不確実性下でも、高い熱的快適性確率と低いオンライン計算時間で安定した性能を示す。

提案手法

  • 状態、行動、報酬関数を備えたマーカフ連鎖過程(MDP)としてHVACシステムをモデル化し、確率的ダイナミクスと熱的快適性制約を捉える。
  • 熱的快適性を定量化するため、非解析的かつ非線形な制約として予測平均投票(PMV)モデルをMDPフレームワークに統合する。
  • ポリシー勾配推定を用いて繰り返し確率的制御ポリシーを更新する勾配ベースポリシー改善(GBPI)アルゴリズムを提案する。
  • ソフトマックス行動選択を用いたパラメータ化された確率的ポリシーを採用し、探索を可能にするとともに、勾配を用いたポリシー最適化を実現する。
  • ポリシー勾配定理に基づいて導出された性能勾配更新ルールを採用することで、最適ポリシーへの収束を保証する。
  • GB-L手法はオフラインで学習され、オンラインで実装され、1秒未塔の応答時間でリアルタイム制御を実現する。

実験結果

リサーチクエスチョン

  • RQ1不確実な熱的擾乱下において、確率的最適制御フレームワークがHVACシステムのエネルギー消費と熱的快適性のバランスを効果的に図れるか。
  • RQ2本手法が、未来の情報を完全に把握できる理想化されたモデル予測制御(MPC-R)と比較して、エネルギーコストと快適性の面でどのように異なるか。
  • RQ3HVACシステムおよびPMVモデルの非凸的かつ非線形な制約が存在する中でも、GB-L手法がグローバル最適ポリシーに収束するか。
  • RQ4本手法におけるエネルギーコスト削減とオンライン計算効率のトレードオフは何か。
  • RQ5変動する天候や占有人数といった現実の不確実性下でも、本手法は高い熱的快適性確率を保証できるか。

主な発見

  • 提案されたGB-L手法は、完全な未来情報が得られる理想的なMPC-Rベースラインと比較して、エネルギーコストがわずか6.5%増加するのみで、近似的最適な性能を達成する。
  • 本手法は1制御ステップあたり1秒未塔のオンライン実行が可能であり、リアルタイム導入に適している。
  • 理論的分析により、与えられたMDP定式化のもとでGB-L手法がグローバル最適ポリシーに収束することを証明した。
  • 不確実性下でも高い熱的快適性確率を確保し、従来の制御戦略に比べて居住者の満足度をより効果的に維持する。
  • シミュレーション結果から、MDPフレームワーク内にPMVモデルを統合することで、エネルギー使用量と熱的快適性のバランスが有効に実現されていることが確認された。
  • 勾配ベースのポリシー更新メカニズムにより、HVACシステムおよびPMVモデルの非凸的かつ非解析的制約を効果的に処理することができた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。