Skip to main content
QUICK REVIEW

[论文解读] A Survey on Popularity Bias in Recommender Systems

Anastasiia Klimashevskaia, Dietmar Jannach|arXiv (Cornell University)|Aug 2, 2023
Recommender Systems and TechniquesComputer Science参考文献 183被引用 3
一句话总结

本综述对推荐系统中的流行度偏差提供了全面分析,探讨了其成因、检测方法及缓解技术。研究指出,当前研究主要基于离线实验,缺乏真实世界验证,呼吁开展更多面向应用场景的评估以及人机协同研究,以提升推荐系统的实际影响与公平性。

ABSTRACT

Recommender systems help people find relevant content in a personalized way. One main promise of such systems is that they are able to increase the visibility of items in the long tail, i.e., the lesser-known items in a catalogue. Existing research, however, suggests that in many situations todays recommendation algorithms instead exhibit a popularity bias, meaning that they often focus on rather popular items in their recommendations. Such a bias may not only lead to the limited value of the recommendations for consumers and providers in the short run, but it may also cause undesired reinforcement effects over time. In this paper, we discuss the potential reasons for popularity bias and review existing approaches to detect, quantify and mitigate popularity bias in recommender systems. Our survey, therefore, includes both an overview of the computational metrics used in the literature as well as a review of the main technical approaches to reduce the bias. Furthermore, we critically discuss todays literature, where we observe that the research is almost entirely based on computational experiments and on certain assumptions regarding the practical effects of including long-tail items in the recommendations.

研究动机与目标

  • 分析推荐系统中流行度偏差的根本原因及其表现形式,特别是其对已有热门项目造成的强化效应。
  • 回顾现有用于检测和量化推荐算法中流行度偏差的计算指标与评估协议。
  • 审视文献中提出的缓解流行度偏差的技术方法,以促进推荐中的多样性与公平性。
  • 批判性评估当前的研究实践,突出对离线实验的过度依赖以及缺乏真实世界验证的问题。
  • 识别研究空白,特别是对面向应用场景的评估、以用户为中心的实验以及长期实地研究的需求,以提升研究的实际相关性。

提出的方法

  • 对推荐系统中流行度偏差的同行评审论文进行系统性文献综述,重点关注定义、评估指标与缓解技术。
  • 将评估指标分类为计算度量,如流行度偏斜、多样性以及公平性指标(例如基尼系数、赫芬达尔-赫希曼指数)。
  • 分析技术缓解策略,包括重排序、损失函数修改以及数据增强,以减少对热门项目的过度依赖。
  • 评估现有实验设置,强调离线基准测试的主导地位以及各研究间缺乏标准化协议的问题。
  • 利用模拟与长期建模研究推荐循环中流行度偏差的长期强化效应。
  • 呼吁方法论多样化,倡导开展人机协同实验与真实世界实地研究,以评估流行度偏差对真实用户的影响。
Figure 1 : Biases and the Feedback Loop of Recommendation, inspired by [ 36 ] .
Figure 1 : Biases and the Feedback Loop of Recommendation, inspired by [ 36 ] .

实验结果

研究问题

  • RQ1推荐系统中流行度偏差的主要来源与定义是什么?不同研究之间有何差异?
  • RQ2哪些计算指标最常用于检测和量化流行度偏差?这些指标在文献中的一致性如何?
  • RQ3现有技术方法在缓解流行度偏差方面的有效性如何?其在真实世界部署中的局限性是什么?
  • RQ4当前评估实践在多大程度上依赖离线实验?这如何影响研究发现的有效性与实际相关性?
  • RQ5为更好地理解流行度偏差对用户、内容提供者及平台生态系统的实际影响,需要哪些方法论上的改进?

主要发现

  • 流行度偏差在推荐系统中普遍存在,算法经常过度推荐已有热门项目,导致多样性下降并强化了‘越富越强’的效应。
  • 尽管文献数量不断增长,但大多数研究仍依赖标准数据集进行离线实验,缺乏真实世界验证或用户感知研究。
  • 使用了多种多样的评估指标,如基尼系数、赫芬达尔-赫希曼指数和流行度偏斜,但尚未就阈值或标准协议达成共识。
  • 各研究之间缺乏可复现性,且数据划分方法不一致,阻碍了不同方法之间的比较与进展。
  • 迫切需要面向应用场景的评估框架,以考虑利益相关者的影响,因为当前指标往往无法反映真实世界的价值或公平性关切。
  • 未来研究必须转向整体性方法,包括人机协同实验与实地研究,以确保偏差缓解策略不仅在计算上有效,而且在实践中具有实际影响力。
Figure 2 : Literature collection methodology.
Figure 2 : Literature collection methodology.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。