[Paper Review] Experimental analysis of data-driven control for a building heating system
This paper proposes a data-driven control approach using model-assisted batch reinforcement learning, specifically fitted Q-iteration with virtual support tuples and domain knowledge shaping, to optimize building heating control under dynamic pricing. It achieves near-optimal performance within 20 days in simulation and demonstrates practical, usable policies in a living lab across varying outdoor temperatures.
Driven by the opportunity to harvest the flexibility related to building climate control for demand response applications, this work presents a data-driven control approach building upon recent advancements in reinforcement learning. More specifically, model assisted batch reinforcement learning is applied to the setting of building climate control subjected to a dynamic pricing. The underlying sequential decision making problem is cast on a markov decision problem, after which the control algorithm is detailed. In this work, fitted Q-iteration is used to construct a policy from a batch of experimental tuples. In those regions of the state space where the experimental sample density is low, virtual support samples are added using an artificial neural network. Finally, the resulting policy is shaped using domain knowledge. The control approach has been evaluated quantitatively using a simulation and qualitatively in a living lab. From the quantitative analysis it has been found that the control approach converges in approximately 20 days to obtain a control policy with a performance within 90% of the mathematical optimum. The experimental analysis confirms that within 10 to 20 days sensible policies are obtained that can be used for different outside temperature regimes.
Motivation & Objective
- To develop a data-driven control strategy for building heating systems that leverages reinforcement learning for demand response applications.
- To address the challenge of sparse data in building climate control by augmenting experimental tuples with virtual samples using a neural network.
- To improve policy performance by integrating domain knowledge into the learned control policy.
- To evaluate the method both quantitatively via simulation and qualitatively in a real-world living lab setting.
- To demonstrate convergence to near-optimal performance under dynamic electricity pricing.
Proposed method
- The control problem is formulated as a Markov decision process (MDP) to model sequential decision-making in building climate control.
- Fitted Q-iteration is employed to learn a control policy from a batch of experimental data tuples collected from the building system.
- In low-data regions of the state space, artificial neural networks generate virtual support tuples to improve policy generalization.
- The resulting policy is refined using domain knowledge to enhance robustness and practicality.
- The approach is validated through simulation and real-world testing in a living lab environment.
- Dynamic electricity pricing is used as the reward signal to incentivize energy cost reduction.
Experimental results
Research questions
- RQ1Can a data-driven control policy be learned efficiently from limited experimental data in a building heating system?
- RQ2How does the inclusion of virtual support tuples via neural networks affect policy performance in low-sampling regions?
- RQ3To what extent can domain knowledge improve the practicality and robustness of a learned reinforcement learning policy?
- RQ4How quickly does the control policy converge to near-optimal performance under dynamic pricing?
- RQ5Can the learned policy generalize across different outdoor temperature regimes in real-world conditions?
Key findings
- The data-driven control approach converges to a policy within 90% of the mathematical optimum in approximately 20 days of training in simulation.
- In the living lab, sensible and usable control policies were obtained within 10 to 20 days, applicable across diverse outdoor temperature conditions.
- The integration of virtual support tuples via neural networks improved policy performance in data-sparse regions of the state space.
- The use of domain knowledge led to more robust and practical control policies compared to raw reinforcement learning output.
- The method demonstrated strong adaptability to dynamic pricing signals, enabling effective demand response in building heating systems.
- The results confirm the feasibility of applying batch reinforcement learning to real-world building climate control with practical convergence times.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.