Skip to main content
QUICK REVIEW

[Paper Review] Penalized EM algorithm and copula skeptic graphical models for inferring networks for mixed variables

Fentaw Abegaz, Ernst C. Wit|arXiv (Cornell University)|Jan 21, 2014
Statistical Methods and Inference22 references3 citations
TL;DR

This paper proposes two novel methods—copula EM glasso and copula skeptic glasso—for inferring sparse graphical models from high-dimensional mixed-type data (continuous, binary, count, and ordinal variables). By combining $\ell_1$-penalized extended rank likelihood with the EM algorithm or pair-wise copula estimation, the methods enable robust network inference under non-normality and mixed variable types, with strong performance in simulation and real-world genomics data, including breast cancer and maize genetics datasets.

ABSTRACT

In this article, we consider the problem of reconstructing networks for continuous, binary, count and discrete ordinal variables by estimating sparse precision matrix in Gaussian copula graphical models. We propose two approaches: $\ell_1$ penalized extended rank likelihood with Monte Carlo Expectation-Maximization algorithm (copula EM glasso) and copula skeptic with pair-wise copula estimation for copula Gaussian graphical models. The proposed approaches help to infer networks arising from nonnormal and mixed variables. We demonstrate the performance of our methods through simulation studies and analysis of breast cancer genomic and clinical data and maize genetics data.

Motivation & Objective

  • To address the challenge of inferring undirected graphical models from high-dimensional mixed-type data (continuous, binary, count, ordinal) where standard Gaussian assumptions fail.
  • To develop a computationally efficient and statistically robust method for estimating sparse precision matrices in Gaussian copula graphical models.
  • To enable network inference that is robust to non-normality and does not require parametric modeling of marginal distributions.
  • To apply the methods to real-world biological datasets, including breast cancer genomic and clinical data and maize genetics data, to uncover biologically relevant conditional dependencies.
  • To compare the performance of the proposed methods with existing approaches using simulation studies and real data analysis, focusing on graph recovery accuracy and computational scalability.

Proposed method

  • Uses the extended rank likelihood (ERL) as a marginal likelihood that is free of nuisance marginal distributions, enabling valid inference for mixed discrete and continuous variables.
  • Applies $\ell_1$-penalized maximum likelihood estimation to the ERL within an EM algorithm framework (copula EM glasso), enabling sparse precision matrix estimation in high dimensions.
  • Employs a copula skeptic approach that estimates pairwise rank correlations via parametric bivariate copulas (not necessarily from the same family), improving accuracy over nonparanormal skeptic methods.
  • Uses the pair-wise copula estimates to construct a correlation matrix, which is then inverted to estimate the precision matrix under $\ell_1$-penalization (copula skeptic glasso).
  • Implements the EM algorithm to handle missing data seamlessly in the copula EM glasso approach, maintaining computational efficiency.
  • Selects tuning parameters via the minimum BIC criterion to balance model fit and sparsity in both methods.

Experimental results

Research questions

  • RQ1Can $\ell_1$-penalized extended rank likelihood with EM algorithm (copula EM glasso) effectively recover the true conditional independence structure in high-dimensional mixed-variable networks?
  • RQ2Does the copula skeptic glasso approach, using parametric bivariate copulas for pairwise rank correlation estimation, outperform existing nonparanormal skeptic methods in terms of graph recovery accuracy?
  • RQ3How do the proposed methods perform in recovering known biological interactions in real-world genomics datasets such as breast cancer and maize genetics data?
  • RQ4What is the relative computational efficiency and scalability of copula EM glasso versus copula skeptic glasso in very high-dimensional settings (e.g., thousands of variables)?
  • RQ5Can the methods detect meaningful trans-acting genetic interactions across chromosomes in maize, despite multiple testing challenges?

Key findings

  • The copula EM glasso method demonstrated strong performance in recovering the true graph structure in simulation studies, particularly in moderately high-dimensional settings with mixed variable types.
  • The copula skeptic glasso method achieved high computational efficiency and was recommended for very high-dimensional data (e.g., thousands of variables), due to its reliance on pairwise copula estimation.
  • In the breast cancer data, the methods identified a sparse network of conditional dependencies between clinical and genetic variables, highlighting genes with amplifications or deletions linked to disease aggressiveness and poor survival outcomes.
  • In the maize genetics dataset, the copula skeptic glasso detected significant inter-chromosomal interactions (e.g., between markers on chromosomes 1–2, 5–8, 6–7), suggesting potential trans-acting regulatory interactions.
  • The minimum BIC criterion selected a tuning parameter of approximately 0.05 for the maize data, indicating a balance between sparsity and biological plausibility.
  • The results support that both methods are robust to non-normality and can effectively model complex dependencies in mixed-type data without requiring parametric modeling of marginal distributions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.