Skip to main content
QUICK REVIEW

[论文解读] Targeting for long-term outcomes

Jeremy J. Yang, Dean Eckles|arXiv (Cornell University)|Oct 29, 2020
Consumer Market Behavior and Pricing参考文献 42被引用 5
一句话总结

本文提出了一种双重稳健的策略学习方法,利用代理索引插补来从短期代理变量估计长期结果,从而在无需等待长期结果的情况下实现最优目标定位。在《波士顿环球报》的两次大规模实地实验中,该方法在三年内实现了400万至500万美元的净正收益,其表现与基于真实长期结果训练的策略相当,且优于使用短期代理变量的策略。

ABSTRACT

Decision makers often want to target interventions so as to maximize an outcome that is observed only in the long-term. This typically requires delaying decisions until the outcome is observed or relying on simple short-term proxies for the long-term outcome. Here we build on the statistical surrogacy and policy learning literatures to impute the missing long-term outcomes and then approximate the optimal targeting policy on the imputed outcomes via a doubly-robust approach. We first show that conditions for the validity of average treatment effect estimation with imputed outcomes are also sufficient for valid policy evaluation and optimization; furthermore, these conditions can be somewhat relaxed for policy optimization. We apply our approach in two large-scale proactive churn management experiments at The Boston Globe by targeting optimal discounts to its digital subscribers with the aim of maximizing long-term revenue. Using the first experiment, we evaluate this approach empirically by comparing the policy learned using imputed outcomes with a policy learned on the ground-truth, long-term outcomes. The performance of these two policies is statistically indistinguishable, and we rule out large losses from relying on surrogates. Our approach also outperforms a policy learned on short-term proxies for the long-term outcome. In a second field experiment, we implement the optimal targeting policy with additional randomized exploration, which allows us to update the optimal policy for future subscribers. Over three years, our approach had a net-positive revenue impact in the range of $4-5 million compared to the status quo.

研究动机与目标

  • 解决主要结果仅在长期才能观测到、从而延迟策略学习的挑战。
  • 开发一种利用统计代理变量插补缺失长期结果的方法,实现及时且基于数据的策略优化。
  • 评估基于插补长期结果学习的策略是否与基于真实长期结果训练的策略表现相当。
  • 通过随机探索实现实时实施与更新最优策略,以应对订阅行为可能存在的非平稳性。
  • 在真实世界订阅管理情境中,特别是大型新闻出版商的主动流失管理中,证明该方法的有效性。

提出的方法

  • 利用代理索引估计器插补长期结果,该估计器基于历史数据中可观测的短期代理变量,建模长期结果的条件期望。
  • 应用双重稳健估计框架,在插补的长期结果上优化目标策略,确保对模型误设的鲁棒性。
  • 使用自助法Thompson抽样以维持探索性,并在多个实验波次中实现动态策略更新。
  • 通过将基于插补结果的策略表现与基于真实长期结果训练的策略及基于短期代理变量的策略进行比较,验证该方法。
  • 采用两阶段实验设计:第一阶段为受控实地实验,用于评估策略表现;第二阶段为后续实验,引入随机探索,以实现持续的策略再优化。
  • 依赖领域知识选择相关代理变量(如短期收入和内容消费),基于其在从处理到长期结果的因果路径中的位置。

实验结果

研究问题

  • RQ1在插补长期结果上优化的策略是否能与在真实长期结果上训练的策略表现相当?
  • RQ2基于代理插补结果的策略表现与基于短期代理变量的策略相比如何?
  • RQ3能否通过随机探索在随时间推移有效更新基于代理插补的策略,以适应非平稳性?
  • RQ4代理插补实现有效策略优化的条件是什么?这些条件与平均处理效应估计的条件有何不同?
  • RQ5在真实世界订阅管理中,基于代理的策略学习在多大程度上能复制基于真实长期结果训练的策略所实现的收益影响?

主要发现

  • 在首次实验中,基于插补长期结果优化的策略在统计上与基于真实长期结果训练的策略表现无显著差异,未检测到性能损失。
  • 基于代理的策略显著优于基于长期收入短期代理变量优化的策略,证明了准确插补相较于简单代理变量的价值。
  • 在三年内,通过随机探索实施优化策略,相较于《波士顿环球报》的现状,实现了400万至500万美元的净正收益。
  • 该方法在策略优化中的有效性成立条件,比平均处理效应估计所需的条件略为宽松,从而具备更广泛的应用潜力。
  • 发现短期收入和内容消费等代理变量是长期留存和收入的有效预测因子,支持其在类似未来应用中的使用。
  • 该方法实现了无需等待长期结果的及时策略学习,使在动态环境中能够实现迭代改进与适应。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。