Skip to main content
QUICK REVIEW

[Paper Review] Using Reinforcement Learning for Demand Response of Domestic Hot Water Buffers: a Real-Life Demonstration

Oscar De Somer, Ana Soares|arXiv (Cornell University)|Mar 16, 2017
Smart Grid Energy Management12 references17 citations
TL;DR

This paper proposes a model-based reinforcement learning (RL) algorithm to optimize domestic hot water (DHW) buffer heating cycles in real residential buildings, maximizing self-consumption of on-site photovoltaic (PV) energy. By learning occupant behavior, predicting PV generation, and accounting for system dynamics, the RL controller increased PV self-consumption by over 20% compared to thermostat control in a six-house field experiment over four months.

ABSTRACT

This paper demonstrates a data-driven control approach for demand response in real-life residential buildings. The objective is to optimally schedule the heating cycles of the Domestic Hot Water (DHW) buffer to maximize the self-consumption of the local photovoltaic (PV) production. A model-based reinforcement learning technique is used to tackle the underlying sequential decision-making problem. The proposed algorithm learns the stochastic occupant behavior, predicts the PV production and takes into account the dynamics of the system. A real-life experiment with six residential buildings is performed using this algorithm. The results show that the self-consumption of the PV production is significantly increased, compared to the default thermostat control.

Motivation & Objective

  • To develop a data-driven control strategy that enables residential demand response using existing thermal storage in domestic hot water (DHW) buffers.
  • To maximize self-consumption of locally generated photovoltaic (PV) electricity in residential buildings without compromising user comfort.
  • To evaluate the real-world performance of a model-based reinforcement learning (RL) approach in a live, multi-house deployment.
  • To assess the impact of RL-based scheduling on energy system flexibility and grid impact in a real-life setting.
  • To compare the performance of the RL controller against conventional thermostat-based control in terms of PV utilization and electricity consumption.

Proposed method

  • A model-based reinforcement learning (RL) framework is used to solve the sequential decision-making problem of scheduling DHW buffer heating cycles.
  • The RL agent learns from real system interactions, modeling stochastic occupant hot water usage and predicting PV generation over time.
  • The state space includes the state-of-charge (SoC) of the DHW buffer, time of day, and predicted PV production; actions are heating to T_min (45°C), T_max (55°C), or delaying heating.
  • A Q-function is trained using a dataset of real system trajectories to approximate the optimal policy for minimizing energy cost and maximizing PV self-consumption.
  • The algorithm uses a receding horizon approach, updating decisions based on real-time sensor data (temperature, flow, electricity use).
  • The system is deployed in six renovated social housing units in the Netherlands, each equipped with a smart heat pump and PV panels.

Experimental results

Research questions

  • RQ1Can a model-based reinforcement learning algorithm effectively schedule DHW buffer heating to maximize self-consumption of on-site PV generation in real residential buildings?
  • RQ2How does RL-based control compare to conventional thermostat control in terms of PV self-consumption and overall electricity usage?
  • RQ3To what extent can occupant behavior and PV generation be predicted and integrated into a real-time control policy for thermal storage?
  • RQ4What is the impact of the RL controller on total household electricity consumption and space heating usage in a real-world deployment?
  • RQ5How robust is the RL policy across different seasons and varying weather and occupancy conditions?

Key findings

  • The RL-based control increased PV self-consumption by 16.94% compared to 6.25% under thermostat control, representing a more than twofold improvement.
  • Total PV self-consumption across all houses rose from 46.69% to 58.52% of total electricity use when using the RL controller.
  • Despite a 10% increase in daily electricity consumption (from 20.17 kWh to 22.13 kWh), the majority of the increase was due to higher space heating use, not DHW.
  • The share of PV captured by space heating increased only slightly, from 18.62% to 19.71%, indicating that the RL policy primarily targeted DHW buffer utilization.
  • The policy effectively shifted DHW heating cycles toward peak PV production times, as shown in the hourly consumption profiles.
  • The results demonstrate that RL can significantly enhance local energy utilization in residential buildings using existing thermal storage infrastructure.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.