Skip to main content
QUICK REVIEW

[论文解读] Flexible Online Repeated Measures Experiment

Guo Yu, Alex Deng|arXiv (Cornell University)|Jan 2, 2015
Advanced Causal Inference Techniques参考文献 11被引用 4
一句话总结

该论文提出了 FORME(灵活的在线重复测量实验),一种可扩展的框架,通过使用重复测量设计(特别是交叉设计和重随机化设计),在不增加流量或实验时长的前提下,提升A/B测试的敏感度。通过利用组内比较,FORME降低了关键绩效指标(KPI)测量中的方差,从而在相同流量和实验时长下实现更高的统计功效,其在缺失数据处理方面的鲁棒性优于传统A/B测试和竞争的混合效应模型。

ABSTRACT

Online controlled experiments, now commonly known as A/B testing, are crucial to causal inference and data driven decision making in many internet based businesses. While a simple comparison between a treatment (the feature under test) and a control (often the current standard), provides a starting point to identify the cause of change in Key Performance Indicator (KPI), it is often insufficient, as the change we wish to detect may be small, and inherent variation contained in data may obscure movements in KPI. To have sufficient power to detect statistically significant changes in KPI, an experiment needs to engage a sufficiently large proportion of traffic to the site, and also last for a sufficiently long duration. This limits the number of candidate variations to be evaluated, and the speed new feature iterations. We introduce more sophisticated experimental designs, specifically the repeated measures design, including the crossover design and related variants, to increase KPI sensitivity with the same traffic size and duration of experiment. In this paper we present FORME (Flexible Online Repeated Measures Experiment), a flexible and scalable framework for these designs. We evaluate the theoretic basis, design considerations, practical guidelines and big data implementation. We compare FORME to an existing methodology called mixed effect model and demonstrate why FORME is more flexible and scalable. We present empirical results based on both simulation and real data. Our method is widely applicable to online experimentation to improve sensitivity in detecting movements in KPI, and increase experimentation capability.

研究动机与目标

  • 解决由于用户层面方差高和处理效应小导致的在线A/B测试统计功效低下的问题。
  • 在不增加流量或实验时长的前提下,提升检测KPI微小变化的敏感度。
  • 开发一种灵活、可扩展的框架,用于在在线实验中实施重复测量设计。
  • 克服现有方法(如混合效应模型)在数据非随机缺失时的局限性。
  • 通过减少达到同等功效所需的样本量,实现更快、更高效的实验周期。

提出的方法

  • FORME采用重复测量设计,包括交叉设计和重随机化设计,使用户在连续的时间段内经历处理组和对照组条件。
  • 通过考虑个体用户的基线水平,利用组内比较降低方差,有效将每个用户作为其自身的对照。
  • 该框架支持灵活的随机化单位,从用户到页面浏览量,实现更精细的处理切换,进一步降低方差。
  • 其统计模型可估计处理效应,同时检验延迟效应和时间趋势,确保在部分周或不规则实验时长下结果的有效性。
  • FORME采用两阶段方法:首先检验处理效应在不同时间是否一致以及是否存在延迟效应,若效应不显著则简化模型。
  • 即使缺失数据与用户层面的随机效应相关,该方法仍能提供稳健估计,其表现优于混合效应模型。

实验结果

研究问题

  • RQ1像交叉设计这样的重复测量设计是否能在不增加流量或实验时长的前提下,提升在线A/B测试的统计功效?
  • RQ2在非忽略性缺失数据条件下,FORME方法与混合效应模型在偏差和方差方面有何比较?
  • RQ3在在线实验中实施重复测量设计的实际设计考量有哪些,特别是关于时间安排和用户体验方面?
  • RQ4通过提高重复测量设计中处理切换的频率,方差可降低到何种程度?
  • RQ5当实验被中断或仅运行部分周时,如何可靠地报告实验结果?

主要发现

  • FORME通过组内比较降低KPI测量的方差,从而实现更高的统计功效,使在相同流量和实验时长下可检测到更小的处理效应。
  • 与传统t检验相比,该框架可将所需样本量减少最多k%(其中k为方差减少因子),该结果基于历史数据中查询量等指标的方差减少观察值。
  • FORME中的交叉设计可将处理效应的方差降低40%-50%,优于标准A/B测试,该结论以先前的CUPED结果作为基准验证。
  • 当数据非随机缺失时,FORME在偏差方面优于混合效应模型,使其在真实在线实验中更具可靠性。
  • 重随机化和多周期设计可检测时间趋势和延迟效应,提升模型有效性,并支持在整个实验周期内报告处理效应。
  • 该方法支持部分周实验而不会造成数据损失,显著提升实用性并减少实验浪费。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。