[论文解读] Causal inference from observational data: Estimating the effect of contributions on visitation frequency atLinkedIn
本文提出并评估了使用观察性 LinkedIn 数据的因果推断方法,以估计用户贡献(如发帖、发消息)对访问频率的即时效应和溢出效应。结果表明,固定效应和双重稳健估计器优于朴素相关性分析,在模型误设情况下也能显著降低偏差;其中公开贡献显示出正向因果效应,而未经适当调整的私密贡献则表现出严重的负向偏差。
Randomized experiments (A/B testings) have become the standard way for web-facing companies to guide innovation, evaluate new products, and prioritize ideas. There are times, however, when running an experiment is too complicated (e.g., we have not built the infrastructure), costly (e.g., the intervention will have a substantial negative impact on revenue), and time-consuming (e.g., the effect may take months to materialize). Even if we can run an experiment, knowing the magnitude of the impact will significantly accelerate the product development life cycle by helping us prioritize tests and determine the appropriate traffic allocation for different treatment groups. In this setting, we should leverage observational data to quickly and cost-efficiently obtain a reliable estimate of the causal effect. Although causal inference from observational data has a long history, its adoption by data scientist in technology companies has been slow. In this paper, we rectify this by providing a brief introduction to the vast field of causal inference with a specific focus on the tools and techniques that data scientist can directly leverage. We illustrate how to apply some of these methodologies to measure the effect of contributions (e.g., post, comment, like or send private messages) on engagement metrics. Evaluating the impact of contributions on engagement through an A/B test requires encouragement design and the development of non-standard experimentation infrastructure, which can consume a tremendous amount of time and financial resources. We present multiple efficient strategies that exploit historical data to accurately estimate the contemporaneous (or instantaneous) causal effect of a user's contribution on her own and her neighbors' (i.e., the users she is connected to) subsequent visitation frequency. We apply these tools to LinkedIn data for several million members.
研究动机与目标
- 使用观察性数据估计用户贡献(公开和私密)对其自身及其社交网络邻居访问频率的因果效应。
- 通过利用历史数据,解决 A/B 测试的局限性,如高成本、基础设施复杂性和延迟效应。
- 在模型误设和未观测混杂因素下,评估不同因果推断方法(尤其是固定效应和双重稳健估计器)的稳健性。
- 开发一个可扩展、安全且标准化的内部平台,用于 LinkedIn 的因果推断,以实现方法论的民主化和严谨性。
- 识别不同用户群体和贡献类型下处理效应的差异,为未来的产品举措提供依据。
提出的方法
- 采用结构模型,将贡献概率和访问频率建模为可观测协变量、时变和时不变未观测混杂因素以及网络邻居行为的函数。
- 应用逆概率加权(IPW)以校正处理分配中的选择偏差。
- 采用双重稳健估计器,结合结果回归模型和倾向得分模型,确保只要任一模型正确设定,估计结果即具一致性。
- 实施固定效应(FE)和加权固定效应模型,以控制用户层面的未观测时不变混杂因素。
- 通过在处理方程和结果方程中引入邻居贡献行为,建模溢出效应,对公开和私密贡献使用对数函数和指示函数。
- 通过已知真实值的模拟验证方法,比较不同混杂情景(时不变 vs. 时变未观测混杂因素)下的偏差和精度。
实验结果
研究问题
- RQ1从观察性数据中估计,用户自身公开或私密贡献对其后续访问频率的因果效应是什么?
- RQ2用户贡献如何影响其社交网络邻居的访问频率(溢出效应)?
- RQ3当存在未观测混杂因素时,不同因果推断方法(IPW、双重稳健、固定效应)在估计这些效应时的表现如何?
- RQ4模型误设(如错误的函数形式或未观测时变混杂因素)如何影响因果估计的准确性?
- RQ5集中化、安全且标准化的平台能否使数据科学家能够可靠地大规模应用先进的因果推断技术?
主要发现
- 贡献与访问频率之间的相关性具有误导性,表现为负相关,而因果方法揭示出正向效应。
- 固定效应(FE)和加权 FE 模型显著降低了由未观测时不变混杂因素引起的偏差,当此类混杂因素存在时,FE 模型几乎完美地估计了真实效应。
- 对于公开贡献,即时效应估计值为 +6.93(时不变)和 +1.34(时变),表明其具有正向但随时间递减的影响。
- 对于私密贡献,时不变效应估计值为 -72.53,但该值严重有偏;真实效应由 FE 模型更准确地捕捉为 -10.00,表明朴素方法高估了其负向影响。
- 双重稳健估计器和 FE 模型在降低偏差方面持续优于标准 IPW 方法,尤其在模型误设情况下表现更优。
- 模拟研究证实,FE 模型对未观测时不变混杂因素具有稳健性,而无固定效应的模型估计值可能比真实值大五倍以上。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。