[论文解读] Using Ego-Clusters to Measure Network Effects at LinkedIn
本文提出了一种名为自我-聚类随机化(ego-cluster randomization)的可扩展、低开销方法,用于测量 LinkedIn 等社交平台中的网络效应,其中用户行为受其社交关系的影响。通过将个体(ego)及其直接联系人(alters)视为聚类,该方法在不改变平台现有系统或引入复杂建模的前提下,保持了统计功效和样本代表性,成功检测到诸如同伴分享带来的参与度提升等下游效应。
A network effect is said to take place when a new feature not only impacts the people who receive it, but also other users of the platform, like their connections or the people who follow them. This very common phenomenon violates the fundamental assumption underpinning nearly all enterprise experimentation systems, the stable unit treatment value assumption (SUTVA). When this assumption is broken, a typical experimentation platform, which relies on Bernoulli randomization for assignment and two-sample t-test for assessment of significance, will not only fail to account for the network effect, but potentially give highly biased results. This paper outlines a simple and scalable solution to measuring network effects, using ego-network randomization, where a cluster is comprised of an "ego" (a focal individual), and her "alters" (the individuals she is immediately connected to). Our approach aims at maintaining representativity of clusters, avoiding strong modeling assumption, and significantly increasing power compared to traditional cluster-based randomization. In particular, it does not require product-specific experiment design, or high levels of investment from engineering teams, and does not require any changes to experimentation and analysis platforms, as it only requires assigning treatment an individual level. Each user either has the feature or does not, and no complex manipulation of interactions between users is needed. It focuses on measuring the one-out network effect (i.e the effect of my immediate connection's treatment on me), and gives reasonable estimates at a very low setup cost, allowing us to run such experiments dozens of times a year.
研究动机与目标
- 解决企业实验中因社交平台网络效应导致的 SUTVA 违反问题。
- 开发一种低成本、可扩展的解决方案,测量‘单向外网络效应’(即连接人对 ego 的影响),而无需产品特定设计或工程开销。
- 在避免强建模假设的前提下,保持统计功效和总体样本的代表性。
- 实现每年数十次的网络效应实验,对现有平台干扰极小。
提出的方法
- 自我-聚类被定义为一个焦点用户(ego)及其直接连接的网络成员(alters),构成自然的随机化聚类单位。
- 处理分配在个体层面进行:每个 ego 被独立地分配为处理组或对照组,不改变平台现有的随机化或分析流水线。
- 通过将用户流失率(例如 20%)限制在较低水平,确保聚类间高度隔离,从而减少外部连接带来的干扰。
- 使用标准的两样本 t 检验评估统计显著性,A/A 测试用于过滤虚假结果。
- 该方法聚焦于测量‘单向外网络效应’——即 ego 的连接人(alters)被处理后,对 ego 行为的影响。
- 通过依赖聚类层面的随机化,避免复杂建模,同时保持总体人群的代表性。
实验结果
研究问题
- RQ1如何在不违反 SUTVA 或不进行重大系统变更的前提下,测量大规模社交平台中的网络效应?
- RQ2自我-聚类随机化能否通过用户连接关系检测到功能对用户行为的有意义下游影响?
- RQ3在优化社交网络参与度时,直接(主效应)与间接(同伴)效应之间的权衡是什么?
- RQ4即使对 alters 的直接效应可忽略不计,细微的内容组成变化在多大程度上能引发可测量的网络效应?
- RQ5将高影响力用户的反馈重新分配给低活跃用户,能否通过网络效应提升整体平台参与度?
主要发现
- 当 ego 的 alters 被分配内容推荐算法(提升分享行为)时,尽管直接参与度下降 5%,自我-聚类随机化仍成功检测到 ego 参与度提升 3.1%(p=0.03)。
- 当 alters 的内容组成被调整以维持相似的分享量时,观察到 ego 的病毒式传播行为增加 7%(p=0.1),表明即使直接指标持平,仍存在间接网络效应。
- 将反馈从高影响力用户重新分配给低反馈用户,导致 ego 内容贡献增加 0.3%(p=0.09)和点赞数增加 1%(p=0.09),表明反馈再分配存在积极的网络效应。
- 该方法在 20% 用户流失率下仍保持统计功效,导致剩余流量的平均参与度和度数下降 15–30%,但结论依然稳健。
- 该方法使每年可执行数十次网络效应实验,工程成本极低,且无需修改现有实验平台。
- A/A 测试证实了结果的可靠性,显著发现被有效过滤,避免了虚假结论。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。