[Paper Review] Multiple imputation of multilevel missing data: An introduction to the R package pan
This paper introduces the R package pan for multiple imputation of multilevel missing data, demonstrating its use via the mitml package to improve accessibility. It shows that accounting for multilevel structure in imputation preserves valid statistical inferences, reducing bias in multilevel model estimates compared to listwise deletion or single-level imputation.
The treatment of missing data can be difficult in multilevel research because state-of-the-art procedures such as multiple imputation (MI) may require advanced statistical knowledge or a high degree of familiarity with certain statistical software. In the missing data literature, pan has been recommended for MI of multilevel data. In this article, we provide an introduction to MI of multilevel missing data using the R package pan, and we discuss its possibilities and limitations in accommodating typical questions in multilevel research. To make pan more accessible to applied researchers, we make use of the mitml package, which provides a user-friendly interface to the pan package and several tools for managing and analyzing multiply imputed data sets. We illustrate the use of pan and mitml with two empirical examples that represent common applications of multilevel models, and we discuss how these procedures may be used in conjunction with other software.
Motivation & Objective
- Address the challenge of applying multiple imputation (MI) to multilevel data, which is often overlooked in applied research.
- Overcome the barrier of technical complexity in using the pan package for multilevel MI, especially for non-expert R users.
- Demonstrate the integration of pan with the mitml package to streamline imputation, diagnostics, and analysis of multiply imputed multilevel datasets.
- Illustrate proper handling of multilevel structures—such as within- and between-group effects and random slopes—during imputation to avoid biased estimates.
- Provide practical guidance on combining results from multiply imputed datasets using Rubin’s rules and alternative methods (D1, D2, D3) for complex hypothesis testing.
Proposed method
- Use the pan package to perform multilevel multiple imputation via a multilevel mixed model (MLMM) that accounts for clustering and random effects.
- Apply the mitml package as a user-friendly interface to pan, enabling structured imputation workflows and automated convergence diagnostics.
- Implement centering of student-level predictors around group means to separate within- and between-group effects in multilevel models.
- Use the imputation model to include auxiliary variables and preserve multilevel structure, ensuring compatibility with the analysis model.
- Apply Rubin’s rules and alternative methods (D1, D2, D3) to pool parameter estimates and test hypotheses across multiple imputed datasets.
- Conduct model comparisons and test constraints (e.g., equality of regression coefficients) using pooling procedures compatible with multiply imputed data.
Experimental results
Research questions
- RQ1How can multiple imputation be effectively applied to multilevel data while preserving the hierarchical structure and avoiding bias in parameter estimates?
- RQ2What are the practical challenges in using the pan package for multilevel MI, and how can the mitml package improve accessibility for applied researchers?
- RQ3How does ignoring the multilevel structure in imputation affect the estimation of within- and between-group effects in multilevel models?
- RQ4What methods are available for pooling results from multiply imputed multilevel datasets when testing complex hypotheses involving multiple parameters?
- RQ5What are the limitations of current multilevel MI approaches when missing data occur in predictor variables of models with random slopes?
Key findings
- Using pan with proper multilevel imputation models significantly reduces bias in multilevel model estimates compared to listwise deletion or single-level imputation.
- The mitml package enables efficient and user-friendly implementation of pan, supporting convergence diagnostics and pooling of results across imputed datasets.
- Imputing multilevel data while preserving group-level structure leads to more accurate estimates of both fixed and random effects, particularly in models with random intercepts and slopes.
- The inclusion of auxiliary variables in the imputation model improves estimation efficiency and helps maintain the integrity of multilevel relationships.
- Pooling procedures such as D1, D2, and D3 allow for valid inference on complex hypotheses, including model comparisons and parameter constraints, across multiply imputed datasets.
- Despite advances, challenges remain in estimating fit indices (e.g., AIC, BIC) and handling missing data in predictors of random-slope models, indicating a need for further methodological development.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.