Skip to main content
QUICK REVIEW

[Paper Review] Potential Impacts of Smart Homes on Human Behavior: A Reinforcement Learning Approach

Shashi Suman, Ali Etemad|arXiv (Cornell University)|Feb 26, 2021
Green IT and Sustainability45 references18 citations
TL;DR

This paper proposes a hierarchical reinforcement learning (HRL) model for simulating human behavior in smart homes, integrated with a Q-learning-based smart home (SHS) that adapts thermal settings for comfort. It finds that SHS can induce unintended behavioral changes—such as frequent activity switching and prolonged time to adjust temperature/humidity—especially when human models have mismatched reward functions, even in multi-agent scenarios.

ABSTRACT

We aim to investigate the potential impacts of smart homes on human behavior. To this end, we simulate a series of human models capable of performing various activities inside a reinforcement learning-based smart home. We then investigate the possibility of human behavior being altered as a result of the smart home and the human model adapting to one-another. We design a semi-Markov decision process human task interleaving model based on hierarchical reinforcement learning that learns to make decisions to either pursue or leave an activity. We then integrate our human model in the smart home which is based on Q-learning. We show that a smart home trained on a generic human model is able to anticipate and learn the thermal preferences of human models with intrinsic rewards similar to the generic model. The hierarchical human model learns to complete each activity and set optimal thermal settings for maximum comfort. With the smart home, the number of time steps required to change the thermal settings are reduced for the human models. Interestingly, we observe that small variations in the human model reward structures can lead to the opposite behavior in the form of unexpected switching between activities which signals changes in human behavior due to the presence of the smart home.

Motivation & Objective

  • To investigate how RL-powered smart homes might unintentionally alter human behavioral patterns in residential environments.
  • To model human agents using hierarchical reinforcement learning (HRL) capable of switching between activities and adjusting thermal preferences.
  • To evaluate the impact of SHS adaptation on human model behavior, particularly in terms of activity switching frequency and time to set thermal comfort.

Proposed method

  • A hierarchical reinforcement learning (HRL) framework models human agents that perform activities (rest, TV, workout) and adjust temperature/humidity for comfort.
  • A Q-learning-based smart home model learns human thermal preferences from feedback, adapting ambient settings to maximize comfort.
  • The human model learns optimal activity duration and transition timing via HRL, while the SHS learns a policy to minimize discomfort through reward maximization.
  • Simulations evaluate single- and multi-human models in a shared home, measuring time-steps to adjust thermal settings and activity switching frequency.
  • Thermal comfort is modeled using PMV (Predicted Mean Vote) within defined ranges, with intrinsic reward functions tailored to individual sensitivity.
  • Experiments compare behavior with and without SHS, analyzing changes in mean time-steps (MTS), mean reward (MR), and PMV trajectory stability.

Experimental results

Research questions

  • RQ1How does the presence of an RL-based smart home affect the activity-switching behavior of a simulated human agent?
  • RQ2To what extent does the SHS alter the time required for a human model to adjust thermal settings (temperature/humidity)?
  • RQ3Can behavioral anomalies such as frequent switching or prolonged adjustment times emerge when the SHS is trained on a different human model than the one it serves?
  • RQ4How do interactions between two human models with differing reward functions and thermal preferences affect behavior in a shared smart home?
  • RQ5Under what conditions do multiple human models cooperate to minimize thermal adjustment time and maximize comfort?

Key findings

  • Human models with reward functions mismatched to the SHS’s training policy exhibited frequent activity switching and increased time-steps (MTS) to adjust thermal settings, rising from 6 to 11 steps in some cases.
  • When two human models shared a home, models with similar thermal preferences (e.g., PMV range [-0.25, 0.25]) showed stable behavior and reduced MTS (from 11 to 5 steps) with SHS, indicating improved coordination.
  • Models with different intrinsic reward functions but overlapping thermal preferences (e.g., HA and HD) reduced MTS to 3–5 steps with SHS, showing improved adaptation and reduced conflict.
  • The SHS learned a stable policy when both models had similar thermal preferences, but bias emerged when one model provided more feedback, affecting policy convergence.
  • In multi-agent scenarios, models adjusted task order to minimize thermal differences, reducing time spent adjusting TH and increasing comfort.
  • PMV trajectory analysis confirmed that SHS improved convergence to comfortable states, especially for sensitive models (e.g., HD), with steeper slopes indicating faster adaptation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.