[Paper Review] Cross-validation Approaches for Multi-study Predictions
This paper proposes a no-data-reuse (NDR) stacking framework for multi-study prediction that combines generalist and specialist prediction models across heterogeneous datasets. By estimating stacking weights without reusing training data, the method achieves oracle-like performance with reduced overfitting, particularly improving accuracy when study-level differences are moderate and sample sizes are large.
We consider prediction in multiple studies with potential differences in the relationships between predictors and outcomes. Our objective is to integrate data from multiple studies to develop prediction models for unseen studies. We propose and investigate two cross-validation approaches applicable to multi-study stacking, an ensemble method that linearly combines study-specific ensemble members to produce generalizable predictions. Among our cross-validation approaches are some that avoid reuse of the same data in both the training and stacking steps, as done in earlier multi-study stacking. We prove that under mild regularity conditions the proposed cross-validation approaches produce stacked prediction functions with oracle properties. We also identify analytically in which scenarios the proposed cross-validation approaches increase prediction accuracy compared to stacking with data reuse. We perform a simulation study to illustrate these results. Finally, we apply our method to predicting mortality from long-term exposure to air pollutants, using collections of datasets.
Motivation & Objective
- Address the challenge of developing robust prediction models when multiple studies have differing distributions of predictors and outcomes.
- Develop a framework that supports both generalist predictions (for new, unseen studies) and specialist predictions (for specific included studies).
- Mitigate overfitting in stacking-based ensemble models caused by data reuse during weight optimization.
- Demonstrate that the proposed no-data-reuse technique yields nearly optimal stacking weights and improves prediction accuracy in finite samples compared to data-reuse methods.
Proposed method
- Use stacking to combine individual study-specific prediction functions (SPFs) into a single ensemble prediction function.
- Estimate stacking weights via a no-data-reuse (NDR) procedure that separates SPF training from utility function estimation.
- Employ task-specific utility functions (e.g., mean squared error) to guide weight optimization without reusing the same data for both training and evaluation.
- Apply asymptotic theory to show that the NDR-stacked model achieves oracle-like performance under mild regularity conditions.
- Use simulation studies to validate theoretical results and compare NDR with traditional data-reuse (DR) stacking in terms of mean squared error.
- Characterize the conditions under which NDR outperforms DR, particularly when study heterogeneity is moderate and sample sizes are large.
Experimental results
Research questions
- RQ1Under what conditions does no-data-reuse stacking outperform data-reuse stacking in multi-study prediction?
- RQ2Can the proposed framework achieve nearly optimal performance comparable to an asymptotic oracle benchmark?
- RQ3How does the accuracy of the stacked prediction model depend on the number of studies and their sample sizes?
- RQ4What is the impact of study heterogeneity (e.g., varying coefficient distributions) on the performance of stacking with and without data reuse?
- RQ5Does the NDR approach reduce overfitting in ensemble models trained on multiple, potentially non-exchangeable studies?
Key findings
- The no-data-reuse (NDR) stacking framework produces stacked prediction functions with nearly optimal performance, approaching the theoretical oracle benchmark under mild regularity conditions.
- When the number of studies $K$ and sample sizes $n_k$ grow large, both NDR and data-reuse (DR) stacking achieve performance close to the asymptotic oracle.
- In a special case with $eta_0 = 0$, the NDR method yields strictly better expected prediction accuracy than DR stacking, with $\mathbb{E}(\psi(\hat{w}^{\text{DR}}) - \psi(\hat{w}^{\text{CS}})) > 0$ for any $K > 2$.
- When study coefficients are close to a common value ($\sigma_\beta \to 0$), the expected difference in prediction error between DR and NDR approaches tends to zero, indicating diminishing gains from NDR.
- As coefficient variability increases ($\sigma_\beta \to \infty$), the expected prediction error of DR stacking diverges relative to NDR, showing that NDR is more robust under high heterogeneity.
- Monte Carlo simulations confirm that $\mathbb{E}(\psi(\hat{w}^{\text{CS}}) - \psi(w_g^0))$ decreases with $K$ as $c_0 + c_1 \log(K)/K$, indicating improved performance with more studies.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.