[论文解读] Counterfactual Inference for Consumer Choice Across Many Product Categories
本文提出了一种可扩展的贝叶斯层次模型,通过利用产品属性和消费者异质性的潜在因子分解,联合估计多个产品类别中的消费者偏好。通过整合随时间变化的价格和缺货事件,该模型在反事实推断方面(尤其是价格敏感度和个性化促销方面)优于孤立的类别模型,在保留数据上的预测准确性和个性化有效性方面均有显著提升。
This paper proposes a method for estimating consumer preferences among discrete choices, where the consumer chooses at most one product in a category, but selects from multiple categories in parallel. The consumer's utility is additive in the different categories. Her preferences about product attributes as well as her price sensitivity vary across products and are in general correlated across products. We build on techniques from the machine learning literature on probabilistic models of matrix factorization, extending the methods to account for time-varying product attributes and products going out of stock. We evaluate the performance of the model using held-out data from weeks with price changes or out of stock products. We show that our model improves over traditional modeling approaches that consider each category in isolation. One source of the improvement is the ability of the model to accurately estimate heterogeneity in preferences (by pooling information across categories); another source of improvement is its ability to estimate the preferences of consumers who have rarely or never made a purchase in a given category in the training data. Using held-out data, we show that our model can accurately distinguish which consumers are most price sensitive to a given product. We consider counterfactuals such as personally targeted price discounts, showing that using a richer model such as the one we propose substantially increases the benefits of personalization in discounts.
研究动机与目标
- 同时对多个产品类别中的消费者需求进行建模,捕捉跨类别的偏好相关性。
- 改进对价格敏感度和异质性的估计,特别是针对低购买频率产品。
- 实现对消费者对个性化折扣和价格变动的反事实响应的准确预测。
- 开发一种利用类别间数据聚合来增强购买历史稀疏消费者的推断能力的模型。
- 基于反事实性能而非标准预测指标调整超参数,以提升实际应用中的效用。
提出的方法
- 使用嵌套因子分解模型,将消费者效用分解为产品属性、价格敏感度和消费者偏好的潜在因子。
- 采用变分推断并结合均值场高斯近似,以实现对大规模数据集的可扩展性。
- 引入“会话”机制,其中价格和可得性保持恒定,从而实现对随时间变化的条件及缺货事件的建模。
- 采用重参数化技巧以优化证据下界(ELBO)的随机梯度下降。
- 引入产品层级和类别层级的潜在因子,分别设置非价格属性和价格敏感度的独立维度。
- 通过在反事实价格变动和缺货事件上的验证来调整超参数,优先考虑反事实性能而非预测准确性。
实验结果
研究问题
- RQ1与孤立的类别模型相比,跨类别联合建模在估计消费者价格敏感度方面有何改进?
- RQ2该模型在估计从未购买或仅少量购买某类商品的消费者偏好方面,准确度如何?
- RQ3随着潜在结构的丰富,模型预测反事实结果(如目标折扣的影响)的能力提升程度如何?
- RQ4整合随时间变化的价格和缺货事件对模型性能有何影响?
- RQ5基于反事实性能调整超参数与基于标准预测指标调整相比,对模型效用的影响如何?
主要发现
- 嵌套因子分解模型实现了-1.7121的中位数自身价格弹性,显著低于多项式对数模型(-1.1841)和嵌套对数模型(-1.2976),表明其估计出更强的价格敏感度。
- 与混合对数模型相比,该模型在产品间平均弹性标准差(SD(Mean) = 1.2008)更低,而混合对数模型为SD(Mean) = 1.7822,表明弹性估计更一致。
- 该模型在预测消费者对价格变动和缺货的响应方面表现出色,保留数据上的反事实准确性显著提升。
- 通过跨类别的信息聚合,该模型能够准确估计从未在某类别购买过的消费者的偏好,从而改善对稀有商品的推断。
- 基于反事实性能的超参数调优,相比仅基于预测质量的调优,显著提升了模型在真实场景中的实用性。
- 在价格变动事件上的验证结果表明,产品属性使用80维潜在因子、价格敏感度使用20维潜在因子,并结合线性价格项的模型为最优选择。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。