[Paper Review] Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement
The paper evaluates two non-experimental causal methods (DML and SPSM) on 663 Facebook ad experiments to see if they can reliably recover ad-caused lifts; neither approach fully succeeds, with DML performing better but still biased.
Despite their popularity, randomized controlled trials (RCTs) are not always available for the purposes of advertising measurement. Non-experimental data is thus required. However, Facebook and other ad platforms use complex and evolving processes to select ads for users. Therefore, successful non-experimental approaches need to "undo" this selection. We analyze 663 large-scale experiments at Facebook to investigate whether this is possible with the data typically logged at large ad platforms. With access to over 5,000 user-level features, these data are richer than what most advertisers or their measurement partners can access. We investigate how accurately two non-experimental methods -- double/debiased machine learning (DML) and stratified propensity score matching (SPSM) -- can recover the experimental effects. Although DML performs better than SPSM, neither method performs well, even using flexible deep learning models to implement the propensity and outcome models. The median RCT lifts are 29%, 18%, and 5% for the upper, middle, and lower funnel outcomes, respectively. Using DML (SPSM), the median lift by funnel is 83% (173%), 58% (176%), and 24% (64%), respectively, indicating significant relative measurement errors. We further characterize the circumstances under which each method performs comparatively better. Overall, despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign's causal effect.
Motivation & Objective
- Assess whether non-experimental data from a large ad platform can recover causal ad effects without RCTs.
- Compare double/debiased machine learning (DML) and stratified propensity score matching (SPSM) in this setting.
- Characterize conditions under which each method performs relatively better or worse.
- Discuss data and platform limitations that hinder reliable causal estimation in online advertising.
Proposed method
- Apply double/debiased machine learning (DML) to estimate causal effects using a rich feature set and cross-validated orthogonalization to reduce regularization bias.
- Evaluate stratified propensity score matching (SPSM) with a deep learning-based propensity model.
- Use an extensive set of campaign- and user-level features to satisfy unconfoundedness assumptions.
- Leverage 663 Facebook ad experiments with large-scale user-impressions data to benchmark against RCTs.
- Report median lifts by funnel and comparative biases between DML and SPSM.
Experimental results
Research questions
- RQ1Can non-experimental methods on platform-logged data sufficiently undo ad-delivery selection to recover causal effects?
- RQ2How do DML and SPSM perform relative to randomized controlled trials in large-scale Facebook experiments?
- RQ3Under what experimental conditions (funnel stage, campaign type) do these methods perform better or worse?
- RQ4What data/logging enhancements would be needed for reliable non-experimental ad measurement?
Key findings
- SPSM performs poorly relative to RCT benchmarks despite extensive features and modeling.
- DML is less upwardly biased than SPSM on average, but residual bias remains substantial.
- Median ad lift from RCTs by funnel: upper 29%, middle 18%, lower 5%.
- Using DML (and SPSM), median lifts by funnel: upper 83% (173%), middle 58% (176%), lower 24% (64%), signaling large relative measurement errors.
- Prospecting campaigns and smaller baseline conversion rates tend to yield comparatively better non-experimental estimates.
- Larger sample sizes, higher test exposure share, and better propensity model performance improve non-experimental estimates, but gaps remain.
- Overall, non-experimental approaches with available data do not reliably estimate causal ad effects; exogenous variation from RCTs remains necessary.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.