Skip to main content
QUICK REVIEW

[Paper Review] High-dimensional consistency in score-based and hybrid structure learning

Preetam Nandy, Alain Hauser|arXiv (Cornell University)|Jul 9, 2015
Bayesian Modeling and Causal Inference4 citations
TL;DR

This paper establishes high-dimensional consistency for score-based and hybrid Bayesian network structure learning methods, particularly Greedy Equivalence Search (GES) and its adaptive variant ARGES, by introducing data-dependent search space restrictions. It proves consistency under sparse high-dimensional settings and demonstrates through simulations that GES and ARGES outperform the PC algorithm in estimation quality while scaling to thousands of variables.

ABSTRACT

Main approaches for learning Bayesian networks can be classified as constraint-based, score-based or hybrid methods. Although high-dimensional consistency results are available for constraint-based methods like the PC algorithm, such results have not been proved for score-based or hybrid methods, and most of the hybrid methods have not even shown to be consistent in the classical setting where the number of variables remains fixed and the sample size tends to infinity. In this paper, we show that consistency of hybrid methods based on greedy equivalence search (GES) can be achieved in the classical setting with adaptive restrictions on the search space that depend on the current state of the algorithm. Moreover, we prove consistency of GES and adaptively restricted GES (ARGES) in several sparse high-dimensional settings. ARGES scales well to sparse graphs with thousands of variables and our simulation study indicates that both GES and ARGES generally outperform the PC algorithm.

Motivation & Objective

  • To address the lack of high-dimensional consistency results for score-based and hybrid structure learning methods like GES and hybrid variants.
  • To establish theoretical consistency of GES and its adaptive variant ARGES in sparse high-dimensional settings where the number of variables grows with sample size.
  • To improve scalability and estimation performance of score-based methods in high-dimensional data by introducing adaptive search space restrictions.
  • To bridge the theoretical gap between constraint-based (e.g., PC) and score-based (e.g., GES) methods by showing their conceptual and algorithmic connections in the multivariate Gaussian setting.
  • To provide practical guidance on penalty parameter selection for (AR)GES using stability selection or extended BIC, enhancing performance in sparse high-dimensional models.

Proposed method

  • Introduces adaptively restricted GES (ARGES), which restricts the search space based on the current state of the algorithm using conditional independence tests.
  • Proves consistency of GES and ARGES in the classical setting (fixed p, n → ∞) under the DAG-perfect assumption, where the true DAG encodes all and only the conditional independence relationships.
  • Establishes high-dimensional consistency of ARGES under assumption (A5), which controls the growth of oracle versions of the algorithm, with sufficient conditions derived via a connection to the Chow-Liu algorithm.
  • Demonstrates that in the multivariate Gaussian setting, GES and PC are closely related: GES uses partial correlation thresholds to add or delete edges, analogous to conditional independence testing.
  • Proposes a new scoring criterion based on rank correlations for nonparanormal distributions, enabling score-based learning beyond Gaussian assumptions.
  • Combines the forward phase of ARGES with the backward phase of selective GES (SGES), leveraging SGES’s polynomial-time complexity and consistency under mild conditions.

Experimental results

Research questions

  • RQ1Can GES and hybrid methods achieve high-dimensional consistency when the number of variables grows with sample size?
  • RQ2What adaptive search space restrictions enable consistent structure learning in high-dimensional sparse settings?
  • RQ3How do score-based methods like GES compare to constraint-based methods like PC in estimation accuracy and scalability?
  • RQ4What is the theoretical connection between GES and constraint-based algorithms like PC in the multivariate Gaussian setting?
  • RQ5Can the consistency of ARGES be extended to learning minimal independence maps under weaker distributional assumptions?

Key findings

  • ARGES achieves high-dimensional consistency under sparse high-dimensional settings when the oracle version of the algorithm satisfies growth condition (A5), which is shown to hold under strong structural assumptions.
  • GES is consistent in the classical setting (fixed p, n → ∞) under the DAG-perfect assumption, even without distributional restrictions beyond the existence of a true DAG encoding all conditional independencies.
  • Simulation results show that both GES and ARGES outperform the PC algorithm in estimation quality, despite PC’s superior scalability and widespread use in high-dimensional applications.
  • The forward phase of ARGES produces an independence map of the true CPDAG in the large-sample limit, enabling consistency when combined with a consistent backward phase like in SGES.
  • A novel connection is established between GES and the Chow-Liu algorithm, providing theoretical insight into the behavior of GES in high-dimensional settings.
  • The paper recommends using stability selection or the extended BIC criterion for penalty parameter selection in (AR)GES, as these outperform standard BIC in sparse high-dimensional models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.