Skip to main content
QUICK REVIEW

[论文解读] Robust Active Preference Elicitation

Phebe Vayanos, Yingxiao Ye|arXiv (Cornell University)|Mar 4, 2020
Economic and Environmental Valuation被引用 6
一句话总结

本文提出了一种基于可调鲁棒优化的稳健主动偏好获取方法,旨在高风险决策场景中,仅使用有限的成对查询时,最大化最坏情况下的效用或最小化最坏情况下的后悔值。该方法分别将离线和在线偏好获取建模为两阶段半和多阶段鲁棒优化问题,其中信息发现依赖于决策。在合成数据和真实世界无家可归者住房分配数据上,该方法在最坏情况后悔值、排名和效用方面均表现出优越性能。

ABSTRACT

We study the problem of eliciting the preferences of a decision-maker through a moderate number of pairwise comparison queries to make them a high quality recommendation for a specific problem. We are motivated by applications in high stakes domains, such as when choosing a policy for allocating scarce resources to satisfy basic needs (e.g., kidneys for transplantation or housing for those experiencing homelessness) where a consequential recommendation needs to be made from the (partially) elicited preferences. We model uncertainty in the preferences as being set based and} investigate two settings: a) an offline elicitation setting, where all queries are made at once, and b) an online elicitation setting, where queries are selected sequentially over time in an adaptive fashion. We propose robust optimization formulations of these problems which integrate the preference elicitation and recommendation phases with aim to either maximize worst-case utility or minimize worst-case regret, and study their complexity. For the offline case, where active preference elicitation takes the form of a two and half stage robust optimization problem with decision-dependent information discovery, we provide an equivalent reformulation in the form of a mixed-binary linear program which we solve via column-and-constraint generation. For the online setting, where active preference learning takes the form of a multi-stage robust optimization problem with decision-dependent information discovery, we propose a conservative solution approach. Numerical studies on synthetic data demonstrate that our methods outperform state-of-the art approaches from the literature in terms of worst-case rank, regret, and utility. We showcase how our methodology can be used to assist a homeless services agency in choosing a policy for allocating scarce housing resources of different types to people experiencing homelessness.

研究动机与目标

  • 解决在仅能获取部分偏好信息的情况下,如何在高风险领域(如稀缺住房或肾脏分配)中做出高质量、稳健推荐的挑战。
  • 将用户偏好的不确定性建模为集合型不确定集,以考虑潜在的不一致、响应错误和模型误设。
  • 开发一个统一框架,将偏好获取与推荐整合为单一鲁棒优化问题,以确保在最坏情况下的性能保障。
  • 研究偏好获取的离线(所有查询预先选定)和在线(自适应、顺序查询选择)两种设置。
  • 确保推荐对对抗性响应具有鲁棒性,并在用户响应不一致或不完整时仍保持强性能。

提出的方法

  • 将离线主动偏好获取问题建模为一个两阶段半的鲁棒优化问题,其中信息发现依赖于决策,且查询选择影响未来的信息可用性。
  • 通过列与约束生成法将离线问题转化为混合整数线性规划,以实现精确求解。
  • 为在线主动偏好获取问题提出一种保守近似方法,将其建模为具有决策依赖信息发现的多阶段鲁棒优化问题。
  • 使用可调鲁棒优化来建模用户效用中的不确定性,假设其在名义效用向量的有界偏差范围内。
  • 引入鲁棒性参数 Γ 以控制最坏情况下的不确定性水平,确保推荐对对抗性响应具有韧性。
  • 将该框架应用于来自无家可归者管理信息系统(HMIS)的真实世界数据,基于机构偏好生成并评估反事实住房分配政策。

实验结果

研究问题

  • RQ1如何设计一种主动偏好获取系统,即使在用户响应不一致或具有对抗性时,也能确保在最坏情况效用或后悔值下的稳健推荐?
  • RQ2将偏好不确定性建模为集合型不确定集,对推荐的质量和鲁棒性有何影响?
  • RQ3所提出的鲁棒优化框架在最坏情况后悔值、排名和效用方面,与最先进方法相比在合成和真实世界数据上的表现如何?
  • RQ4该查询选择方法的鲁棒性在多大程度上对模型误设敏感,例如对响应不一致性标准差的错误假设?
  • RQ5所提出的框架能否有效应用于真实世界高风险政策问题,如将稀缺住房资源分配给无家可归者?

主要发现

  • 所提方法在合成和真实世界数据上,均持续优于最先进方法,在最坏情况后悔值、最坏情况真实后悔值和最坏情况真实排名方面表现更优。
  • 仅经过 7 次成对查询,即使在响应不一致(σ = 0.025)的情况下,方法仍能实现最坏情况排名 6 或更优;在无响应不一致(σ = 0)时,能将最佳项目排在首位。
  • 该方法对模型误设表现出强鲁棒性:当 σ ≠ σ* 时,最坏情况后悔值和真实后悔值分别仅下降 0.03 和 0.03(离线)以及 0.00 和 0.01(在线)。
  • 在模型误设下,最坏情况真实排名平均仅下降 2.44(离线)和 0.48(在线),表明具有高度韧性。
  • 在线保守近似方法在实践中实现了接近最优的性能,与精确离线解相比损失极小。
  • 该框架成功识别出一种在公平性、效率和可解释性之间取得平衡的住房分配政策,其依据是来自无家可归者服务机构的偏好信息。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。