Skip to main content
QUICK REVIEW

[论文解读] Potential Impacts of Smart Homes on Human Behavior: A Reinforcement Learning Approach

Shashi Suman, Ali Etemad|arXiv (Cornell University)|Feb 26, 2021
Green IT and Sustainability参考文献 45被引用 18
一句话总结

本文提出一种分层强化学习(HRL)模型,用于模拟智能家居中的人类行为,该模型与基于Q-learning的智能家居系统(SHS)相结合,后者可自适应调节热舒适设置。研究发现,即使在多智能体场景下,当人类模型的奖励函数与SHS不匹配时,SHS仍可能引发非预期的行为变化,如频繁切换活动、延长调节温湿度的时间。

ABSTRACT

We aim to investigate the potential impacts of smart homes on human behavior. To this end, we simulate a series of human models capable of performing various activities inside a reinforcement learning-based smart home. We then investigate the possibility of human behavior being altered as a result of the smart home and the human model adapting to one-another. We design a semi-Markov decision process human task interleaving model based on hierarchical reinforcement learning that learns to make decisions to either pursue or leave an activity. We then integrate our human model in the smart home which is based on Q-learning. We show that a smart home trained on a generic human model is able to anticipate and learn the thermal preferences of human models with intrinsic rewards similar to the generic model. The hierarchical human model learns to complete each activity and set optimal thermal settings for maximum comfort. With the smart home, the number of time steps required to change the thermal settings are reduced for the human models. Interestingly, we observe that small variations in the human model reward structures can lead to the opposite behavior in the form of unexpected switching between activities which signals changes in human behavior due to the presence of the smart home.

研究动机与目标

  • 探究基于强化学习的智能家居在住宅环境中是否可能无意中改变人类行为模式。
  • 使用分层强化学习(HRL)建模人类智能体,使其能够切换活动并调节热舒适偏好。
  • 评估SHS自适应调节对人类模型行为的影响,特别是活动切换频率和调节热舒适所需时间。

提出的方法

  • 采用分层强化学习(HRL)框架建模人类智能体,使其执行活动(如休息、看电视、锻炼)并调节温湿度以实现舒适。
  • 基于Q-learning的智能家居系统(SHS)通过反馈学习人类热舒适偏好,自适应调整环境设置以最大化舒适度。
  • 人类模型通过HRL学习最优活动持续时间和状态转换时机,而SHS则通过奖励最大化学习最小化不适的策略。
  • 通过仿真评估单人与多人模型在共享家庭环境中的表现,测量调节热舒适设置所需的时间步数及活动切换频率。
  • 热舒适度使用PMV(预测平均投票)模型进行建模,范围在预设区间内,且为个体敏感度量身定制内在奖励函数。
  • 实验对比有无SHS情况下的行为表现,分析平均时间步数(MTS)、平均奖励(MR)及PMV轨迹稳定性。

实验结果

研究问题

  • RQ1基于强化学习的智能家居存在时,对模拟人类智能体的活动切换行为有何影响?
  • RQ2SHS在多大程度上改变了人类模型调节热舒适设置(温度/湿度)所需的时间?
  • RQ3当SHS训练所用的人类模型与实际服务对象不一致时,是否会出现如频繁切换或调节时间延长等行为异常?
  • RQ4当两个具有不同奖励函数和热舒适偏好的人类模型共处一个智能家居时,其行为如何相互影响?
  • RQ5在何种条件下,多个人类模型能够协作以最小化热调节时间并最大化舒适度?

主要发现

  • 奖励函数与SHS训练策略不匹配的人类模型表现出频繁的活动切换行为,且调节热舒适设置所需的时间步数(MTS)显著上升,某些情况下从6步增至11步。
  • 当两个人类模型共享一个家庭时,若其热舒适偏好相近(如PMV范围[-0.25, 0.25]),在SHS作用下行为趋于稳定,MTS从11步降至5步,表明协调性提升。
  • 即使内在奖励函数不同但热舒适偏好重叠的模型(如HA与HD)在SHS作用下MTS降至3–5步,表明适应性增强且冲突减少。
  • 当两个模型热舒适偏好相似时,SHS能学习到稳定策略;但若其中一模型提供更多信息反馈,则策略收敛过程可能出现偏差。
  • 在多智能体场景中,模型通过调整任务顺序以最小化热差异,从而减少调节温湿度的时间,提升舒适度。
  • PMV轨迹分析表明,SHS显著提升了向舒适状态的收敛速度,尤其对敏感模型(如HD)更为明显,更陡峭的斜率表明适应速度更快。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。