[Paper Review] A Two-Part Machine Learning Approach to Characterizing Network Interference in A/B Testing
This paper proposes a two-part machine learning approach to detect and characterize heterogeneous network interference in A/B testing by introducing 'causal network motifs' and using transparent models to automate exposure mapping. It outperforms traditional methods like cluster randomization and neighborhood exposure mapping in simulations and real-world tests with 1–2 million Instagram users.
The reliability of controlled experiments, commonly referred to as "A/B tests," is often compromised by network interference, where the outcomes of individual units are influenced by interactions with others. Significant challenges in this domain include the lack of accounting for complex social network structures and the difficulty in suitably characterizing network interference. To address these challenges, we propose a machine learning-based method. We introduce "causal network motifs" and utilize transparent machine learning models to characterize network interference patterns underlying an A/B test on networks. Our method's performance has been demonstrated through simulations on both a synthetic experiment and a large-scale test on Instagram. Our experiments show that our approach outperforms conventional methods such as design-based cluster randomization and conventional analysis-based neighborhood exposure mapping. Our approach provides a comprehensive and automated solution to address network interference for A/B testing practitioners. This aids in informing strategic business decisions in areas such as marketing effectiveness and product customization.
Motivation & Objective
- To address the challenge of network interference in A/B testing, which undermines the validity of causal inference by violating the Stable Unit Treatment Value Assumption (SUTVA).
- To overcome the limitations of existing methods—particularly manual or rigid exposure mapping and reliance on strong parametric assumptions—by automating the identification of interference patterns.
- To develop a scalable, interpretable framework that characterizes heterogeneous interference across complex network structures in large-scale digital platforms.
- To enable retrospective analysis of past A/B tests and guide future experimental design by revealing underlying interference mechanisms.
Proposed method
- Introduces 'causal network motifs' as interpretable, structural features representing recurring interference patterns in networks, enabling feature engineering for machine learning models.
- Applies transparent, interpretable machine learning models (e.g., decision trees or generalized linear models) to learn exposure mapping functions from observed treatment and outcome data.
- Uses a two-stage framework: first, mining causal network motifs from network topology; second, training ML models on these motifs to predict exposure conditions and estimate treatment effects.
- Employs simulation-based validation on synthetic networks (Watts-Strogatz and Slashdot) and real-world data from a large-scale Instagram A/B test involving 1–2 million users.
- Validates performance against conventional methods such as design-based cluster randomization and analysis-based neighborhood exposure mapping.
- Incorporates model interpretability to allow practitioners to understand and act on the identified interference patterns, supporting strategic decisions in product and marketing.
Experimental results
Research questions
- RQ1How can we automatically identify and characterize heterogeneous network interference patterns in A/B testing without relying on pre-specified exposure mappings?
- RQ2To what extent can causal network motifs improve the accuracy of treatment effect estimation in the presence of complex interference?
- RQ3Can a machine learning-based approach outperform traditional design-based and analysis-based methods in estimating global average treatment effects under interference?
- RQ4How do interference patterns vary across different network structures, and can the method generalize across diverse network topologies?
Key findings
- The proposed method significantly outperformed conventional approaches such as design-based cluster randomization and analysis-based neighborhood exposure mapping in both synthetic and real-world experiments.
- In a large-scale A/B test with 1–2 million Instagram users, the method demonstrated superior precision in estimating treatment effects, reducing bias caused by network interference.
- The use of causal network motifs enabled the identification of non-trivial interference patterns that were missed by standard exposure mapping techniques.
- The method achieved robust performance across different network types, including Watts-Strogatz and Slashdot networks, confirming its generalizability.
- The transparent machine learning models provided interpretable exposure mappings, allowing practitioners to understand and refine intervention strategies.
- The approach enabled retrospective analysis of past experiments, revealing hidden interference effects and informing better design for future A/B tests.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.