[论文解读] Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching Approach
本文提出了一种受限匹配框架,通过确保内容提供商维持最低参与度阈值,以优化推荐系统中的长期社会福利。通过将推荐策略建模为最优受限匹配问题,该方法在不显式优化公平性或提供商效用的情况下,显著提升了用户福利和提供商多样性,优于短视策略。
Most recommender systems (RS) research assumes that a user's utility can be maximized independently of the utility of the other agents (e.g., other users, content providers). In realistic settings, this is often not true---the dynamics of an RS ecosystem couple the long-term utility of all agents. In this work, we explore settings in which content providers cannot remain viable unless they receive a certain level of user engagement. We formulate the recommendation problem in this setting as one of equilibrium selection in the induced dynamical system, and show that it can be solved as an optimal constrained matching problem. Our model ensures the system reaches an equilibrium with maximal social welfare supported by a sufficiently diverse set of viable providers. We demonstrate that even in a simple, stylized dynamical RS model, the standard myopic approach to recommendation---always matching a user to the best provider---performs poorly. We develop several scalable techniques to solve the matching problem, and also draw connections to various notions of user regret and fairness, arguing that these outcomes are fairer in a utilitarian sense.
研究动机与目标
- 解决推荐系统研究中长期用户与内容提供商之间生态系统动态被忽视的空白。
- 将用户效用与提供商可行性之间的互动建模为一个动力系统,其中提供商需达到最低参与度才能保持活跃。
- 开发一种策略优化框架,以在确保多样化可行提供商的前提下最大化长期社会福利。
- 证明整体匹配策略在用户福利和提供商多样性方面优于短视推荐。
提出的方法
- 将推荐问题形式化为最优受限匹配问题,以选择社会福利最大的均衡。
- 为提供商引入可行性阈值,低于该阈值时提供商将退出平台,以模拟现实中的激励机制。
- 使用离散事件动力系统模拟随时间推移的用户查询与提供商匹配过程,效用来源于累积的内容曝光。
- 开发可扩展算法以计算最优受限匹配策略,利用潜在嵌入和亲和度得分。
- 在优化目标中引入折现因子,以平衡平均用户福利与最大用户遗憾。
- 建立所生成策略与功利主义公平性的联系,表明在未显式施加公平性约束的情况下,实现了更低的遗憾和更高的多样性。
实验结果
研究问题
- RQ1一种短视推荐策略(始终选择亲和度最高的提供商)在动态生态系统中对长期用户福利和提供商多样性有何影响?
- RQ2一种考虑提供商可行性阈值的受限匹配方法,是否能带来高于标准短视策略的整体社会福利?
- RQ3优化长期用户效用在多大程度上能隐式提升提供商多样性并减少用户遗憾?
- RQ4目标函数中不同水平的折现如何在平均福利与最大遗憾之间进行权衡?
- RQ5所提出策略的结果在何种意义上与功利主义意义上的公平性相一致?
主要发现
- 在折现因子为 1.0 时,所提出的 LP-RS 策略实现了 63.62(±3.58)的平均用户福利和 34.20 个可行提供商,而短视基线仅实现 48.59 的福利和 11.80 个可行提供商。
- 在折现因子为 0.59 时,LP-RS 实现了 34.30 的平均福利和 43.00 个可行提供商,而短视策略仅实现 26.07 的福利和 11.80 个可行提供商。
- LP-RS 下的最大遗憾始终低于短视策略,峰值分别为 γ=1.0 时的 37.89 和短视策略下的 36.62,表明对个体用户的保护更优。
- 尽管目标函数仅优化用户效用,受限匹配方法仍隐式维持了远高于短视策略的可行提供商数量。
- 基于亲和度的随机采样表现劣于 LP-RS 和短视策略,表明仅靠随机化无法维持提供商可行性。
- 福利-遗憾权衡曲线表明,LP-RS 在保持最大遗憾相对较低的同时,维持了较高的社会福利,尤其在收益递减效用函数下表现更优。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。