[Paper Review] Practical targeted learning from large data sets by survey sampling
This paper proposes a practical targeted learning approach for large data sets using survey sampling with unequal inclusion probabilities to reduce computational burden. By selecting a sub-sample based on summary statistics and applying targeted minimum loss estimation (TMLE), the method achieves asymptotically normal estimators and valid confidence intervals, with inclusion probabilities optimized to minimize variance.
We address the practical construction of asymptotic confidence intervals for smooth (i.e., path-wise differentiable), real-valued statistical parameters by targeted learning from independent and identically distributed data in contexts where sample size is so large that it poses computational challenges. We observe some summary measure of all data and select a sub-sample from the complete data set by Poisson rejective sampling with unequal inclusion probabilities based on the summary measures. Targeted learning is carried out from the easier to handle sub-sample. We derive a central limit theorem for the targeted minimum loss estimator (TMLE) which enables the construction of the confidence intervals. The inclusion probabilities can be optimized to reduce the asymptotic variance of the TMLE. We illustrate the procedure with two examples where the parameters of interest are variable importance measures of an exposure (binary or continuous) on an outcome. We also conduct a simulation study and comment on its results. keywords: semiparametric inference; survey sampling; targeted minimum loss estimation (TMLE)
Motivation & Objective
- Address computational challenges in targeted learning when sample size N is extremely large.
- Develop a practical method to construct asymptotic confidence intervals for smooth, pathwise differentiable statistical parameters in large-scale data settings.
- Enable efficient inference by leveraging survey sampling techniques—specifically rejective sampling with unequal inclusion probabilities—on summary statistics of the full data.
- Optimize inclusion probabilities to minimize the asymptotic variance of the targeted minimum loss estimator (TMLE).
- Demonstrate the method’s validity and efficiency through theoretical results and simulation studies on variable importance measures.
Proposed method
- Use Poisson rejective sampling with unequal inclusion probabilities based on summary measures of the full data set to select a manageable sub-sample of size n(N) << N.
- Apply targeted minimum loss estimation (TMLE) to the sub-sample to estimate the parameter of interest, ensuring double robustness and efficiency.
- Derive a central limit theorem for the TMLE under the sampling design, enabling construction of asymptotic confidence intervals.
- Optimize inclusion probabilities using the influence function to minimize the asymptotic variance of the TMLE estimator.
- Leverage the Horvitz-Thompson empirical measure and empirical process theory to establish asymptotic normality and uniform convergence under the sampling scheme.
- Use functional central limit theorems and entropy conditions to validate the asymptotic distribution of the estimator under weak regularity assumptions.
Experimental results
Research questions
- RQ1Can targeted learning be practically applied to very large data sets where full-data computation is infeasible?
- RQ2How can survey sampling with unequal inclusion probabilities improve the efficiency of TMLE in large-scale settings?
- RQ3What is the asymptotic distribution of the TMLE under rejective sampling with unequal probabilities?
- RQ4Can inclusion probabilities be optimized to reduce the asymptotic variance of the TMLE in large data contexts?
- RQ5How does the proposed method perform in estimating variable importance measures for exposure-outcome relationships in large data sets?
Key findings
- The proposed method enables valid asymptotic confidence intervals for smooth statistical parameters using only a sub-sample of the full data set.
- The central limit theorem for the TMLE holds under the rejective sampling design, ensuring valid inference even when N is very large.
- Inclusion probabilities can be optimized using the influence function to minimize the asymptotic variance of the TMLE estimator.
- The method achieves asymptotic normality of the TMLE under weak regularity conditions, including entropy and moment constraints on the model space.
- Simulation studies show that the method maintains good coverage and efficiency, even when the full data set is too large for standard computation.
- The approach is particularly effective for estimating variable importance measures in exposure-outcome models, with demonstrated robustness and computational feasibility.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.