[论文解读] Cluster Randomized Designs for One-Sided Bipartite Experiments
本文提出了一种新颖的聚类随机化设计,用于单边二部图实验,其中干扰通过干扰单元(如市场中的卖家)这一中介层产生。该方法引入了一种平衡划分聚类目标,以在具有有界干扰的线性潜在结果模型下最小化差异均值估计器的偏差,证明了其极小极大最优性,并在多种图结构和模型下展现出鲁棒性。
The conclusions of randomized controlled trials may be biased when the outcome of one unit depends on the treatment status of other units, a problem known as interference. In this work, we study interference in the setting of one-sided bipartite experiments in which the experimental units - where treatments are randomized and outcomes are measured - do not interact directly. Instead, their interactions are mediated through their connections to interference units on the other side of the graph. Examples of this type of interference are common in marketplaces and two-sided platforms. The cluster-randomized design is a popular method to mitigate interference when the graph is known, but it has not been well-studied in the one-sided bipartite experiment setting. In this work, we formalize a natural model for interference in one-sided bipartite experiments using the exposure mapping framework. We first exhibit settings under which existing cluster-randomized designs fail to properly mitigate interference under this model. We then show that minimizing the bias of the difference-in-means estimator under our model results in a balanced partitioning clustering objective with a natural interpretation. We further prove that our design is minimax optimal over the class of linear potential outcomes models with bounded interference. We conclude by providing theoretical and experimental evidence of the robustness of our design to a variety of interference graphs and potential outcomes models.
研究动机与目标
- 解决单边二部图实验中的干扰问题,其中实验单元通过干扰单元这一中介层受到处理暴露的影响。
- 使用暴露映射框架形式化该设置下的干扰,识别现有聚类随机化设计的局限性。
- 在具有有界干扰的线性潜在结果模型下,开发一种最小化差异均值估计器偏差的聚类目标。
- 证明所提出的方案在具有有界干扰的线性潜在结果模型类上的极小极大最优性。
- 通过理论和实证方法验证其对多样化干扰图结构和潜在结果模型的鲁棒性。
提出的方法
- 使用暴露映射框架形式化单边二部图实验中的干扰,建模干扰单元如何在实验单元之间中介处理溢出。
- 提出一种基于实验单元平衡划分的新型聚类目标,该目标源自在具有线性潜在结果的模型下最小化差异均值估计器偏差。
- 定义一种新颖的聚类目标,考虑干扰单元的结构及其在中介中的作用,采用基于边权重(由共享干扰单元决定)的图结构公式。
- 通过平衡划分算法实现聚类,其中实验单元的节点权重设为1,边权重基于共享干扰单元的连接关系。
- 对 EXPOSURE-DESIGN 基线使用贪心局部搜索算法,超参数 λ ∈ {0, 0.001, 0.01, 0.1, 1} 进行调优,通过正则化提升聚类平衡性。
- 应用差异均值估计器来估计处理效应,并通过在多种合成和现实世界图结构下的相对均方根误差(RMSE)与标准差来评估性能。
实验结果
研究问题
- RQ1现有聚类随机化设计能否在不引入偏差的情况下直接应用于单边二部图实验?
- RQ2在具有有界干扰的线性潜在结果模型下,何种聚类目标能最小化差异均值估计器的偏差?
- RQ3所提出的聚类设计是否在具有有界干扰的线性潜在结果模型类上达到极小极大最优?
- RQ4该设计对干扰图结构和潜在结果模型参数的变化具有多大鲁棒性?
- RQ5在估计精度和方差方面,该设计是否优于标准单元级随机化及其他聚类基线?
主要发现
- 所提出的平衡划分聚类目标在具有有界干扰的线性潜在结果模型下,能最小化差异均值估计器的偏差。
- 该设计被证明在具有有界干扰的线性潜在结果模型类上达到极小极大最优,确保对最坏情况干扰配置的鲁棒性。
- 在实验中,该设计在真实聚类结构下实现了 0.31(±0.03) 的相对 RMSE,显著优于 EXPOSURE-DESIGN(0.37±0.04)和单元级随机化(0.39±0.05)。
- 随着邻域宽度 Δ 增加,该设计保持了较低的相对 RMSE(Δ=0.3 时为 0.458±0.005)和标准差(0.025±0.003),优于直接聚类和单元级随机化。
- 该设计对非线性表现出鲁棒性,当 Δ=0.5 时,相对 RMSE 降至 0.001±0.0,表明在纯暴露条件下偏差极小。
- 采用平衡划分算法实现,设置 10 个聚类和 10% 的不平衡容忍度,该设计在多种图结构和模型配置下均表现出一致的性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。