[Paper Review] multilevLCA: An R Package for Single-Level and Multilevel Latent Class Analysis with Covariates
This paper introduces multilevLCA, an R package that enables simultaneous and two-step estimation for single-level and multilevel latent class analysis with covariates, supporting both fixed- and random-effect models. It offers semi-automatic model selection, class membership prediction, and visualization tools, filling a critical gap in R-based LCA software by integrating advanced multilevel modeling with user-friendly, efficient estimation for categorical data in social science research.
This contribution presents a guide to the R package multilevLCA, which offers a complete and innovative set of technical tools for the latent class analysis of single-level and multilevel categorical data. We describe the available model specifications, mainly falling within the fixed-effect or random-effect approaches. Maximum likelihood estimation of the model parameters, enhanced by a refined initialization strategy, is implemented either simultaneously, i.e., in one-step, or by means of the more advantageous two-step estimator. The package features i) semi-automatic model selection when a priori information on the number of classes is lacking, ii) predictors of class membership, and iii) output visualization tools for any of the available model specifications. All functionalities are illustrated by means of a real application on citizenship norms data, which are available in the package.
Motivation & Objective
- To address the lack of comprehensive, user-friendly R packages for multilevel latent class analysis with covariates.
- To implement both simultaneous and two-step estimation strategies for improved algorithmic stability and convergence speed.
- To provide semi-automatic model selection when the number of latent classes is unknown a priori.
- To enable prediction of class membership and visualization of response probabilities across classes.
- To support both fixed-effect and random-effect approaches for multilevel data structures, particularly in survey and cross-national research.
Proposed method
- Uses maximum likelihood estimation with a refined initialization strategy for improved convergence in latent class models.
- Implements a two-step estimator that separates measurement model estimation from structural model estimation, enhancing stability and speed.
- Supports both fixed-effect and random-effect multilevel modeling, with the latter allowing flexible, nonparametric modeling of group-level random effects.
- Integrates semi-automatic model selection via BIC, AIC, and ICL criteria across multiple class and group-level configurations.
- Provides class-profile visualization tools with customizable labels for interpretable presentation of response probabilities.
- Enables inclusion of covariates at both individual and group levels to model predictors of class membership.

Experimental results
Research questions
- RQ1How can multilevel latent class models with covariates be efficiently estimated in R, particularly when the number of classes is unknown?
- RQ2What are the relative advantages of simultaneous versus two-step estimation in multilevel LCA, especially in terms of bias, efficiency, and convergence?
- RQ3How can model selection be automated in the absence of prior knowledge about the number of latent classes?
- RQ4To what extent can class membership be predicted and visualized in multilevel LCA with covariates?
- RQ5How does multilevLCA compare to existing R packages in supporting complex, multilevel LCA with covariates?
Key findings
- The two-step estimator in multilevLCA provides approximately unbiased estimates with minimal efficiency loss compared to one-step maximum likelihood, while offering superior algorithmic stability and faster convergence.
- The optimal model for the citizenship norms data was selected as 3 low-level classes and 2 high-level classes (iT=3, iM=2), with a BIC of 881,252.15 and ICL-BIC of 936,851.08.
- The selected model identified three distinct adolescent citizenship profiles: 'Maximal' (25% probability), 'Engaged' (75% probability), and 'Subject' (0% probability) for the first individual.
- The package successfully visualized response probabilities across classes, with custom labels such as 'Maximal', 'Engaged', and 'Subject' improving interpretability.
- multilevLCA outperforms existing R packages like poLCA, glca, and stepmixr by supporting multilevel models with covariates, both estimation types, and model selection in a single, integrated framework.
- The package enables bias-corrected three-step estimation via minor coding, extending its utility for advanced methodological applications such as differential item functioning and direct effect analysis.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.