Skip to main content
QUICK REVIEW

[论文解读] To Model or to Intervene: A Comparison of Counterfactual and Online Learning to Rank from User Interactions

Rolf Jagerman, Harrie Oosterhuis|arXiv (Cornell University)|Jul 15, 2019
Mobile Crowdsensing and Crowdsourcing参考文献 40被引用 6
一句话总结

本文首次基于用户交互数据,对反事实学习排序(counterfactual)与在线学习排序(online LTR)方法进行了直接比较。研究发现,在低偏差、低噪声环境下,反事实 LTR 表现优于在线 LTR;但在高选择偏差、位置偏差或交互噪声环境下,尽管初始用户体验有所下降,在线 LTR 表现更稳健且更有效。

ABSTRACT

Learning to Rank (LTR) from user interactions is challenging as user feedback often contains high levels of bias and noise. At the moment, two methodologies for dealing with bias prevail in the field of LTR: counterfactual methods that learn from historical data and model user behavior to deal with biases; and online methods that perform interventions to deal with bias but use no explicit user models. For practitioners the decision between either methodology is very important because of its direct impact on end users. Nevertheless, there has never been a direct comparison between these two approaches to unbiased LTR. In this study we provide the first benchmarking of both counterfactual and online LTR methods under different experimental conditions. Our results show that the choice between the methodologies is consequential and depends on the presence of selection bias, and the degree of position bias and interaction noise. In settings with little bias or noise counterfactual methods can obtain the highest ranking performance; however, in other circumstances their optimization can be detrimental to the user experience. Conversely, online methods are very robust to bias and noise but require control over the displayed rankings. Our findings confirm and contradict existing expectations on the impact of model-based and intervention-based methods in LTR, and allow practitioners to make an informed decision between the two methodologies.

研究动机与目标

  • 为解决在真实世界用户交互场景下,反事实与在线学习排序(LTR)方法之间缺乏直接实证比较的问题。
  • 探究选择偏差、位置偏差和交互噪声如何影响反事实与在线 LTR 方法的性能。
  • 根据系统的偏差与噪声特征,为实践者提供在反事实与在线 LTR 之间进行选择的指导。
  • 评估两种方法在不同实验条件下对鲁棒性与用户体验的影响。
  • 为实践者提供基于模型(反事实)与基于干预(在线)的 LTR 方法选择基准。

提出的方法

  • 使用标准基准数据集和评估指标,对反事实 LTR(CLTR)与在线 LTR(OLTR)方法进行全面实证比较。
  • 采用受控实验设置,在不同水平的选择偏差、位置偏差和交互噪声下,对 CLTR 与 OLTR 方法进行评估。
  • 以生产环境排序器作为基线,比较在不同偏差/噪声条件下模型的性能表现。
  • 应用成熟的 CLTR 方法,依赖精确倾向性得分进行去偏处理,确保在理想条件下实现无偏估计。
  • 采用 OLTR 方法(如 PDGD,策略梯度下降),通过主动干预调整排序结果以探索并优化性能。
  • 使用离线指标评估两种方法,并追踪用户体验随时间的变化,特别关注在线 LTR 初期体验下降与长期收益。

实验结果

研究问题

  • RQ1在何种条件下,反事实 LTR 在排序性能上优于在线 LTR?
  • RQ2选择偏差的存在如何影响反事实与在线 LTR 方法的相对性能?
  • RQ3交互噪声对反事实与在线 LTR 方法的性能与鲁棒性有何影响?
  • RQ4位置偏差水平如何影响每种方法的有效性?
  • RQ5在线 LTR 初期对用户体验的损害程度如何?其恢复或超越生产排序器的速度如何?

主要发现

  • 在选择偏差低、位置偏差低且交互噪声最小的环境下,反事实 LTR 表现最佳。
  • 在高选择偏差、位置偏差或交互噪声存在时,在线 LTR 始终优于反事实 LTR。
  • 即使理论上无偏,当交互噪声较高时,反事实 LTR 的性能仍可能劣于生产排序器。
  • 在线 LTR 方法初期用户体验劣于生产排序器,但迅速改善并超越,展现出显著的长期收益。
  • 反事实 LTR 中模型优化与部署的循环可能导致性能下降,原因在于噪声效应的累积。
  • 在线 LTR 比反事实 LTR 更能抵御各类偏差与噪声,因此在现实、嘈杂的生产环境中更具优势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。