Skip to main content
QUICK REVIEW

[Paper Review] Identification and quantification of Granger causality between gene sets

André Fujita, João Ricardo Sato|ArXiv.org|Nov 6, 2009
Bioinformatics and Genomic Networks52 references3 citations
TL;DR

This paper introduces a novel method to identify and quantify Granger causality between sets of genes using canonical correlation analysis (CCA) and bootstrap hypothesis testing, enabling the detection of information flow between biological pathways. The approach outperforms standard VAR models in simulated and real gene expression data, offering improved power to detect causal relationships in high-dimensional biological networks.

ABSTRACT

Wiener and Granger have introduced an intuitive concept of causality between two variables which is based on the idea that an effect never occurs before its cause. Later, Geweke has generalized this concept to a multivariate Granger causality, i.e., n variables Granger-cause another variable. Although Granger causality is not "effective causality", this concept is useful to infer directionality and information flow in observational data. Granger causality is usually identified by using VAR models due to their simplicity. In the last few years, several VAR-based models were presented in order to model gene regulatory networks. Here, we generalize the multivariate Granger causality concept in order to identify Granger causalities between sets of gene expressions, i.e., whether a set of n genes Granger-causes another set of m genes, aiming at identifying and quantifying the flow of information between gene networks (or pathways). The concept of Granger causality for sets of variables is presented. Moreover, a method for its identification with a bootstrap test is proposed. This method is applied in simulated and also in actual biological gene expression data in order to model regulatory networks. This concept may be useful to understand the complete information flow from one network or pathway to the other, mainly in regulatory networks. Linking this concept to graph theory, sink and source can be generalized to node sets. Moreover, hub and centrality for sets of genes can be defined based on total information flow. Another application is in annotation, when the functionality of a set of genes is unknown, but this set is Granger caused by another set of genes which is well studied. Therefore, this information may be useful to infer or construct some hypothesis about the unknown set of genes.

Motivation & Objective

  • To extend Granger causality from individual genes to sets of genes, enabling analysis of information flow between biological pathways or networks.
  • To develop a statistically robust method for identifying and quantifying causal influences between gene sets in high-dimensional gene expression data.
  • To overcome limitations of standard VAR models in detecting multivariate causalities when the number of genes exceeds the number of samples.
  • To provide a framework for inferring functional roles of uncharacterized gene sets based on their causal relationships with well-studied pathways.

Proposed method

  • The method uses canonical correlation analysis (CCA) to model the relationship between two sets of time series (gene sets), capturing linear dependencies between past values of one set and present values of another.
  • A bootstrap-based hypothesis test is employed to assess the statistical significance of the CCA-derived Granger causality, controlling type I error at 5%.
  • The approach quantifies total information flow from one gene set to another by aggregating causal influences across individual genes in the sets.
  • The method is compared against standard VAR models using Wald’s test for multivariate Granger causality, with performance evaluated on simulated and real gene expression data.
  • The framework allows generalization of network concepts like source, sink, hub, and centrality to sets of genes based on total information flow.

Experimental results

Research questions

  • RQ1Can Granger causality be meaningfully extended from individual genes to sets of genes to model information flow between biological pathways?
  • RQ2How can statistical significance of set-level Granger causality be reliably assessed in high-dimensional, low-sample-size gene expression data?
  • RQ3Does the proposed CCA-based method outperform standard VAR models in detecting multivariate Granger causalities between gene sets?
  • RQ4Can the method infer functional roles of poorly characterized gene sets based on their causal relationships with well-annotated pathways?

Key findings

  • The CCA-based method demonstrated higher statistical power than standard VAR models in detecting Granger causalities between gene sets in simulated data with 10,000 replicates.
  • The bootstrap-based testing procedure effectively controlled the false positive rate at 5% across multiple simulations, ensuring robust inference.
  • In simulated data with known causal structures (e.g., I → II, I → III), the method correctly identified causal relationships with high sensitivity and specificity.
  • The method successfully identified biologically plausible information flow patterns in real gene expression data, including time-delayed regulatory influences between pathways.
  • The framework enables the definition of network centrality measures for gene sets, identifying key regulatory hubs based on total information flow.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.