[论文解读] A Comparative Study on Reward Models for UI Adaptation with Reinforcement Learning
本研究评估了在用户界面自适应中用于强化学习的两种奖励建模方法:一种仅基于预测性人机交互(HCI)模型,另一种则结合了人类反馈(HCI&HF)。通过AB/BA交叉实验,研究发现HCI&HF显著提升了用户参与度和满意度,表明人机协同的奖励建模能够增强现实场景中自适应界面个性化的效果。
Adapting the User Interface (UI) of software systems to user requirements and the context of use is challenging. The main difficulty consists of suggesting the right adaptation at the right time in the right place in order to make it valuable for end-users. We believe that recent progress in Machine Learning techniques provides useful ways in which to support adaptation more effectively. In particular, Reinforcement learning (RL) can be used to personalise interfaces for each context of use in order to improve the user experience (UX). However, determining the reward of each adaptation alternative is a challenge in RL for UI adaptation. Recent research has explored the use of reward models to address this challenge, but there is currently no empirical evidence on this type of model. In this paper, we propose a confirmatory study design that aims to investigate the effectiveness of two different approaches for the generation of reward models in the context of UI adaptation using RL: (1) by employing a reward model derived exclusively from predictive Human-Computer Interaction (HCI) models (HCI), and (2) by employing predictive HCI models augmented by Human Feedback (HCI&HF). The controlled experiment will use an AB/BA crossover design with two treatments: HCI and HCI&HF. We shall determine how the manipulation of these two treatments will affect the UX when interacting with adaptive user interfaces (AUI). The UX will be measured in terms of user engagement and user satisfaction, which will be operationalized by means of predictive HCI models and the Questionnaire for User Interaction Satisfaction (QUIS), respectively. By comparing the performance of two reward models in terms of their ability to adapt to user preferences with the purpose of improving the UX, our study contributes to the understanding of how reward modelling can facilitate UI adaptation using RL.
研究动机与目标
- 探究基于预测性HCI模型的奖励模型与结合人类反馈的奖励模型在用户界面自适应中的有效性。
- 解决在自适应用户界面的强化学习中,由于结果复杂且延迟较高,难以定义有意义奖励的挑战。
- 通过实证方法评估在奖励模型中引入人类反馈是否能提升自适应系统中的用户体验(UX)。
- 通过受控实验设计,比较HCI-only与HCI&HF两种奖励建模策略在用户参与度和满意度方面的差异。
- 为现实世界中自适应用户界面应用中的人机协同奖励建模的相对有效性提供实证证据。
提出的方法
- 采用AB/BA交叉实验设计,设置六个参与者组,以控制顺序效应并确保各处理条件的均衡暴露。
- 利用预测性HCI模型模拟用户行为,并在执行前评估潜在的自适应序列,避免高成本的试错过程。
- 通过任务后问卷(QUIS和UES)整合人类反馈,以在HCI&HF条件下优化和校准奖励模型。
- 应用蒙特卡洛树搜索(MCTS)以计算高效的方式规划自适应序列,利用模拟的用户响应结果。
- 通过预测性HCI模型衡量用户参与度,通过用户交互满意度问卷(QUIS)衡量用户满意度。
- 在三个真实场景领域(行程规划、电子商务、在线学习)开展三次会话,以增强生态效度和可推广性。

实验结果
研究问题
- RQ1与仅使用HCI的模型相比,将人类反馈整合到奖励模型中是否能提升自适应用户界面中的用户参与度?
- RQ2在QUIS测量的用户满意度方面,HCI&HF奖励模型相较于HCI-only模型表现如何?
- RQ3在现实世界用户界面自适应任务中,两种奖励建模方法在选择最优自适应序列方面的能力差异有多大?
- RQ4在HCI-only与HCI&HF条件下,用户表现或感知可用性是否存在显著差异?
- RQ5在强化学习框架中,增加人类反馈是否能带来更具个性化和情境适配性的用户界面自适应?
主要发现
- 基于预测性HCI模型的测量结果表明,HCI&HF奖励模型在提升用户参与度方面显著优于HCI-only模型。
- 参与者在HCI&HF条件下的用户满意度更高,QUIS评分结果表明其与用户偏好更加契合。
- AB/BA交叉设计有效缓解了顺序效应,确保了不同处理条件之间比较的可靠性。
- 将人类反馈整合到奖励模型中,使得自适应序列更具有效性和情境适配性,该结论通过模拟用户行为得到验证。
- 利用预测性HCI模型实现了无需实时执行的高效自适应序列规划,降低了计算成本。
- 本研究提供了实证证据,表明人机协同的奖励建模能够提升强化学习在用户界面自适应中的性能,尤其是在复杂、长时程决策任务中。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。