[论文解读] Revenue Maximization and Learning in Products Ranking
本文提出了一种在线零售商的收益最大化框架,将消费者注意力持续时间建模为随机变量,扩展了级联模型以纳入随机注意力和收益目标。该框架设计了一种近似算法,在已知注意力持续时间分布下可实现最优收益的 $1/e$;同时提出一种在线学习算法,即使在购买数据被右删失的情况下,也能实现 $ ilde{/mathcal{O}}(ackslash sqrt{T})$ 的遗憾,数值实验验证了其有效性。
We consider the revenue maximization problem for an online retailer who plans to display in order a set of products differing in their prices and qualities. Consumers have attention spans, i.e., the maximum number of products they are willing to view, and inspect the products sequentially before purchasing a product or leaving the platform empty-handed when the attention span gets exhausted. Our framework extends the well-known cascade model in two directions: the consumers have random attention spans instead of fixed ones, and the firm maximizes revenues instead of clicking probabilities. We show a nested structure of the optimal product ranking as a function of the attention span when the attention span is fixed. \sg{Using this fact, we develop an approximation algorithm when only the distribution of the attention spans is given. Under mild conditions, it achieves $1/e$ of the revenue of the clairvoyant case when the realized attention span is known. We also show that no algorithms can achieve more than 0.5 of the revenue of the same benchmark. The model and the algorithm can be generalized to the ranking problem when consumers make multiple purchases.} When the conditional purchase probabilities are not known and may depend on consumer and product features, we devise an online learning algorithm that achieves $ ilde{\mathcal{O}}(\sqrt{T})$ regret relative to the approximation algorithm, despite the censoring of information: the attention span of a customer who purchases an item is not observable. Numerical experiments demonstrate the outstanding performance of the approximation and online learning algorithms.
研究动机与目标
- 将消费者在产品排序中的行为建模为随机注意力持续时间,突破经典级联模型中固定注意力持续时间的限制。
- 在注意力持续时间分布和产品吸引力存在不确定性的情况下,解决收益最大化问题。
- 设计一种近似算法,在已知注意力持续时间分布时,可实现克拉沃伊ant基准收益的 $1/e$。
- 开发一种在线学习算法,能够适应未知的产品和消费者特征,最小化遗憾,尽管存在非购买用户导致的删失数据。
- 将框架推广至多商品购买场景,并通过数值实验验证性能。
提出的方法
- 通过从已知分布中抽取随机注意力持续时间,扩展级联模型,允许消费者在注意力耗尽时离开而不购买。
- 推导出固定注意力持续时间下的最优排序的嵌套结构,从而通过动态规划或贪心选择实现高效近似。
- 提出一种近似算法,在温和条件下可保证实现克拉沃伊ant情况(即已知实际注意力持续时间)下收益的 $1/e$。
- 设计一种基于上下文Bandit的在线学习算法,遗憾为 $ ilde{ackslash mathcal{O}}(ackslash sqrt{T})$,能够处理购买决策隐藏真实注意力持续时间的删失反馈。
- 通过在注意力约束下建模序列决策,将算法应用于多商品购买场景。
- 结合理论分析与数值实验,验证近似算法和学习算法的性能。
实验结果
研究问题
- RQ1当注意力持续时间随机且目标为收益最大化而非点击率最大化时,最优产品排序是什么?
- RQ2当仅知道注意力持续时间分布时,如何近似最优排序?此类近似的性能保证是什么?
- RQ3能否设计一种在线学习算法,以适应未知的产品和消费者特征,同时在存在删失数据的情况下最小化遗憾?
- RQ4在此设置下,任何算法的性能基本极限是什么?所提方法与该界限相比如何?
- RQ5该模型如何推广到消费者单次会话中进行多次购买的场景?
主要发现
- 对于固定注意力持续时间,最优排序表现出嵌套结构,即 $k$ 个商品的排序是 $k+1$ 个商品排序的子集,从而支持高效计算。
- 在温和条件下,所提近似算法可实现克拉沃伊ant情况(即真实注意力持续时间已知)下收益的 $1/e \approx 0.368$。
- 任何算法都无法实现超过克拉沃伊ant基准收益的 $0.5$,从而确立了理论上的上界。
- 尽管存在非购买用户导致的删失反馈,所提在线学习算法相对于近似算法的遗憾为 $\tilde{ackslash mathcal{O}}(ackslash sqrt{T})$。
- 数值实验表明,近似算法和在线学习算法在各种设置下均表现出强劲的实证性能。
- 该框架可自然推广至多商品购买场景,在保持理论保证的同时具备实际有效性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。