[论文解读] User Response Prediction in Online Advertising
本综述系统地梳理了在线广告中用户响应预测的机器学习与深度学习方法,涵盖平台、利益相关方、数据类型和技术方法。综述了最先进模型、基准数据集及开源实现,以指导点击率预测与用户行为建模领域的工业界与学术界研究。
Online advertising, as the vast market, has gained significant attention in various platforms ranging from search engines, third-party websites, social media, and mobile apps. The prosperity of online campaigns is a challenge in online marketing and is usually evaluated by user response through different metrics, such as clicks on advertisement (ad) creatives, subscriptions to products, purchases of items, or explicit user feedback through online surveys. Recent years have witnessed a significant increase in the number of studies using computational approaches, including machine learning methods, for user response prediction. However, existing literature mainly focuses on algorithmic-driven designs to solve specific challenges, and no comprehensive review exists to answer many important questions. What are the parties involved in the online digital advertising eco-systems? What type of data are available for user response prediction? How to predict user response in a reliable and/or transparent way? In this survey, we provide a comprehensive review of user response prediction in online advertising and related recommender applications. Our essential goal is to provide a thorough understanding of online advertising platforms, stakeholders, data availability, and typical ways of user response prediction. We propose a taxonomy to categorize state-of-the-art user response prediction methods, primarily focus on the current progress of machine learning methods used in different online platforms. In addition, we also review applications of user response prediction, benchmark datasets, and open-source codes in the field.
研究动机与目标
- 为在线广告生态系统提供系统性理解,包括广告商、发布商与平台等关键利益相关方。
- 识别并分类不同广告平台中用于用户响应预测的数据与特征类型。
- 构建用于预测用户响应(如点击、转化、购买)的最先进机器学习与深度学习方法的结构化分类体系。
- 综述实际工业应用、基准数据集与开源实现,以支持可复现性与进一步研究。
- 解决用户响应预测系统在真实部署中面临的可靠性、可解释性与可扩展性等关键挑战。
提出的方法
- 提出一种分类体系,基于模型架构对用户响应预测方法进行分类,包括因子分解机、深度神经网络、图神经网络与混合模型。
- 综述代表性模型如DeepFM、DIN、DIEN、PIN与DCN_V2,重点分析其在学习用户-物品交互与用户兴趣演化方面的机制。
- 分析工业级系统如Facebook的FBCTR、Etsy的EtsyCTR与Google的DCN_V2,聚焦特征哈希、模型并行与分布式训练等技术在可扩展性方面的应用。
- 考察基于图的模型如PinSage与RippleNet,分析其利用用户-物品交互图提升表征学习的效果。
- 引入结合协同过滤、基于内容的特征与注意力机制的混合方法,以建模动态用户偏好。
- 评估处理长用户行为序列的技术,包括记忆网络(MIMN、SIM)与自注意力模块,用于行为检索与相关性评分。

实验结果
研究问题
- RQ1在线广告生态系统中的关键组成部分与利益相关方有哪些?它们在用户响应预测中如何互动?
- RQ2在不同广告平台(如搜索广告、展示广告、应用内广告)中,可用于建模用户响应的数据与特征类型有哪些?
- RQ3现代机器学习与深度学习模型如何提升点击率与转化率预测的准确性与效率?
- RQ4在大规模部署用户响应预测模型时面临的主要技术挑战是什么?工业系统如何应对?
- RQ5与传统方法相比,混合模型与基于图的模型在用户兴趣建模与长期行为理解方面有何增强作用?
主要发现
- 综述指出,用户响应预测最常被建模为点击率(CTR)估计问题,深度学习模型在准确性与适应性方面显著优于传统方法。
- 图神经网络(GNN)如PinSage与RippleNet通过建模复杂的用户-物品交互图提升性能,尤其在推荐与展示广告中表现突出。
- 工业系统如Facebook的FBCTR与Google的DCN_V2通过子采样、模型并行与参数共享等技术实现高可扩展性,支持实时推理。
- DIEN与SIM等模型引入注意力机制与记忆网络,以捕捉用户兴趣的演化与长序列行为,提升对动态用户画像的预测能力。
- 混合模型如DeepFM与PIN结合因子分解机与深度神经网络,有效学习低阶与高阶特征交互,在基准数据集上达到最先进性能。
- 综述强调,开源实现与基准数据集(如Criteo、Avazu、MovieLens)是推动可复现性与领域进步的关键,PyTorch与Caffe2等工具在模型部署中被广泛使用。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。