Skip to main content
QUICK REVIEW

[论文解读] A/B Testing: A Systematic Literature Review

Federico Quin, Danny Weyns|arXiv (Cornell University)|Aug 9, 2023
Software Engineering Research被引用 6
一句话总结

本篇系统性文献回顾分析了141篇关于A/B测试的原始研究,以绘制A/B测试研究的最新进展。研究识别出关键目标(算法、视觉元素、工作流程)、主导的测试类型(采用假设检验的经典A/B测试)、利益相关者角色(设计师、架构师、技术人员、协调员、评估员)、收集的数据类型(产品、以用户为中心、时空数据),以及结果的常见用途(功能选择与发布)。本研究提出三个未来研究方向:改进统计方法、优化A/B测试流程,以及提升自动化水平。

ABSTRACT

In A/B testing two variants of a piece of software are compared in the field from an end user's point of view, enabling data-driven decision making. While widely used in practice, no comprehensive study has been conducted on the state-of-the-art in A/B testing. This paper reports the results of a systematic literature review that analyzed 141 primary studies. The results shows that the main targets of A/B testing are algorithms and visual elements. Single classic A/B tests are the dominating type of tests. Stakeholders have three main roles in the design of A/B tests: concept designer, experiment architect, and setup technician. The primary types of data collected during the execution of A/B tests are product/system data and user-centric data. The dominating use of the test results are feature selection, feature rollout, and continued feature development. Stakeholders have two main roles during A/B test execution: experiment coordinator and experiment assessor. The main reported open problems are enhancement of proposed approaches and their usability. Interesting lines for future research include: strengthen the adoption of statistical methods in A/B testing, improving the process of A/B testing, and enhancing the automation of A/B testing.

研究动机与目标

  • 提供一个全面且基于实证的当前A/B测试研究现状概述。
  • 从研究视角识别A/B测试的主要目标、设计模式、执行实践及利益相关者角色。
  • 揭示A/B测试中的开放挑战与研究空白,以指导未来学术与工业界的研究。
  • 通过识别改进A/B测试流程与工具的机会,支持实践者。

提出的方法

  • 在主要计算机科学数字图书馆中使用预定义的搜索字符串开展系统性文献回顾。
  • 应用纳入与排除标准,筛选出141篇原始研究,排除简短论文、演示文稿或路线图类论文。
  • 采用质量评分系统(≤4分者被排除)以确保研究的可靠性并减少偏倚。
  • 对A/B测试目标、设计、执行、利益相关者角色、数据类型及结果使用情况进行数据提取。
  • 通过滚雪球法与交叉核对提升外部效度并减少检索偏倚。
  • 基于发现的题域分析,对开放问题进行分类,并推导出未来研究方向。

实验结果

研究问题

  • RQ1A/B测试在软件系统中的主要目标是什么?
  • RQ2A/B测试在实践中如何设计与执行,利益相关者扮演何种角色?
  • RQ3A/B测试过程中收集了哪些类型的数据,结果如何被使用?
  • RQ4文献中报告的主要开放挑战与局限性是什么?
  • RQ5A/B测试中最具前景的未来研究方向是什么?

主要发现

  • 最常见的A/B测试目标是算法、视觉元素以及工作流程或流程变更,其主要应用领域为网络、搜索引擎和电子商务。
  • 采用两个变体的经典A/B测试占主导地位,主要依赖假设检验进行结果分析,而自助法(bootstrapping)在部分研究中逐渐受到关注。
  • 利益相关者在设计阶段扮演三种不同角色:概念设计师、实验架构师与设置技术人员;在执行阶段扮演两种角色:实验协调员与实验评估员。
  • 最常收集的数据类型为产品/系统数据、以用户为中心的数据以及时空数据,支持超越基础指标的深入分析。
  • 测试结果主要用于功能选择、渐进式功能发布、持续开发,以及为后续A/B测试设计提供依据。
  • 识别出七类开放问题,包括改进现有方法、提升评估与分析能力,以及增强A/B测试流程的自动化与可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。