[论文解读] Trustworthy Online Marketplace Experimentation with Budget-split Design
本文提出了预算拆分设计(budget-split design),一种针对双边在线市场的新型实验框架,通过将活动预算拆分至处理组与对照组,消除竞争效应偏差(cannibalization bias)。该设计实现无偏估计并显著提升统计功效——在检测能力方面,相较其他设计仅能实现5–12%的统计功效,本方法可达到80%的检测能力,从而能够识别出以往因统计功效不足而被遗漏的有意义产品影响。
Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient power. In two-sided online marketplaces, both requirements are called into question. The Bernoulli randomized experiments are biased because treatment units interfere with control units through market competition and violate the "stable unit treatment value assumption"(SUTVA). The experimental power on at least one side of the market is often insufficient because of disparate sample sizes on the two sides. Despite the important of online marketplaces to the online economy and the crucial role experimentation plays in product improvement, there lacks an effective and practical solution to the bias and low power problems in marketplace experimentation. Our paper fills this gap by proposing an experimental design that is unbiased in any marketplace where buyers have a defined budget, which could be finite or infinite. We show that it is more powerful than all other unbiased designs in literature. We then provide generalizable system architecture for deploying this design to online marketplaces. Finally, we confirm our findings with empirical performance from experiments run in two real-world online marketplaces.
研究动机与目标
- 解决在线市场实验中因市场竞争和样本量不平衡导致的偏差与低统计功效问题。
- 开发一种实用且假设极少的实验设计,确保在双边市场中实现无偏估计。
- 基于真实用户和活动数据,在真实在线市场中实证验证该设计。
- 证明预算拆分实验在偏差控制与统计功效方面,优于活动级别、轮换(switchback)及聚类设计。
提出的方法
- 预算拆分设计为每项活动分配一个总预算,并将其拆分为处理组与对照组部分,确保各单元之间无干扰。
- 该设计基于潜在结果框架,定义了对目标结果平均处理效应的估计量(estimand)。
- 在预算可拆分的市场单侧存在的前提下,推导出处理效应的无偏估计量,仅依赖于该前提。
- 通过可扩展的系统架构实现部署,支持大规模、生产级的实验应用。
- 实验结果采用两样本t检验进行分析,并与活动级别及轮换设计的统计功效进行比较。
- 通过将预算拆分实验结果与已知存在偏差的成员级别实验结果对比,验证其对竞争效应偏差的鲁棒性。
实验结果
研究问题
- RQ1如何设计一种既无偏又具高统计功效的双边在线市场实验框架?
- RQ2与传统伯努利随机化实验相比,预算拆分设计在多大程度上减少了偏差?
- RQ3在真实市场中,预算拆分实验的统计功效与活动级别及轮换实验相比如何?
- RQ4预算拆分设计能否检测到低功效替代方案所遗漏的有意义处理效应?
- RQ5当未考虑干扰效应时,标准市场实验中的竞争效应偏差程度有多大?
主要发现
- 预算拆分实验对处理效应的检测统计功效达到80%,而活动级别实验在Marketplace 1中仅为5.2%(Marketplace 2为12%)。
- 轮换实验的统计功效分别为5.1%和5.2%,仅略高于5%的显著性水平,实际效果可忽略不计。
- 预算拆分设计成功消除了竞争效应偏差,否则该偏差会导致成员级别实验中处理效应被高估100–200%。
- 该设计成功检测到价值数百万美元的年度收入损失,这些影响在活动级别实验中因统计功效不足而长期未被发现。
- 实证结果证实,预算拆分实验对干扰具有鲁棒性,且不具有成员级别或活动级别设计固有的偏差。
- 该方法仅需市场单侧存在可拆分预算这一基本假设,无需额外建模假设,因此具有广泛适用性,适用于真实环境的部署。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。