Skip to main content
QUICK REVIEW

[论文解读] Identification and quantification of Granger causality between gene sets

André Fujita, João Ricardo Sato|ArXiv.org|Nov 6, 2009
Bioinformatics and Genomic Networks参考文献 52被引用 3
一句话总结

本文提出了一种新颖的方法,利用典型相关分析(CCA)和自助法假设检验,识别并量化基因集之间的格兰杰因果关系,从而实现对生物通路间信息流的检测。该方法在模拟数据和真实基因表达数据中均优于标准VAR模型,在高维生物网络中检测因果关系的统计功效更优。

ABSTRACT

Wiener and Granger have introduced an intuitive concept of causality between two variables which is based on the idea that an effect never occurs before its cause. Later, Geweke has generalized this concept to a multivariate Granger causality, i.e., n variables Granger-cause another variable. Although Granger causality is not "effective causality", this concept is useful to infer directionality and information flow in observational data. Granger causality is usually identified by using VAR models due to their simplicity. In the last few years, several VAR-based models were presented in order to model gene regulatory networks. Here, we generalize the multivariate Granger causality concept in order to identify Granger causalities between sets of gene expressions, i.e., whether a set of n genes Granger-causes another set of m genes, aiming at identifying and quantifying the flow of information between gene networks (or pathways). The concept of Granger causality for sets of variables is presented. Moreover, a method for its identification with a bootstrap test is proposed. This method is applied in simulated and also in actual biological gene expression data in order to model regulatory networks. This concept may be useful to understand the complete information flow from one network or pathway to the other, mainly in regulatory networks. Linking this concept to graph theory, sink and source can be generalized to node sets. Moreover, hub and centrality for sets of genes can be defined based on total information flow. Another application is in annotation, when the functionality of a set of genes is unknown, but this set is Granger caused by another set of genes which is well studied. Therefore, this information may be useful to infer or construct some hypothesis about the unknown set of genes.

研究动机与目标

  • 将格兰杰因果关系从单个基因扩展至基因集,以实现对生物通路或网络间信息流的分析。
  • 开发一种统计上稳健的方法,用于识别和量化高维基因表达数据中基因集之间的因果影响。
  • 克服标准VAR模型在基因数超过样本数时检测多变量因果关系的局限性。
  • 提供一种框架,基于基因集与已知通路的因果关系,推断未充分表征基因集的功能角色。

提出的方法

  • 该方法使用典型相关分析(CCA)建模两组时间序列(基因集)之间的关系,捕捉一组基因集过去值与另一组基因集当前值之间的线性依赖关系。
  • 采用基于自助法的假设检验,评估由CCA推导出的格兰杰因果关系的统计显著性,控制第一类错误率在5%。
  • 通过聚合基因集中各基因的因果影响,量化从一个基因集到另一个基因集的总信息流。
  • 将该方法与标准VAR模型进行比较,采用Wald检验评估多变量格兰杰因果关系,性能在模拟和真实基因表达数据上进行评估。
  • 该框架可将网络概念如源、汇、枢纽和中心性推广至基因集,基于总信息流进行定义。

实验结果

研究问题

  • RQ1格兰杰因果关系能否有意义地从单个基因扩展至基因集,以建模生物通路间的信息流?
  • RQ2在高维、小样本量的基因表达数据中,如何可靠地评估基因集层面格兰杰因果关系的统计显著性?
  • RQ3所提出的基于CCA的方法是否在检测基因集之间的多变量格兰杰因果关系方面优于标准VAR模型?
  • RQ4该方法能否基于基因集与已注释通路之间的因果关系,推断出功能表征不足的基因集的功能角色?

主要发现

  • 在10,000次重复的模拟数据中,基于CCA的方法在检测基因集之间格兰杰因果关系时,统计功效高于标准VAR模型。
  • 基于自助法的检验程序在多次模拟中有效控制了假阳性率在5%,确保了推断的稳健性。
  • 在具有已知因果结构的模拟数据中(如 I → II,I → III),该方法以高敏感性和特异性正确识别了因果关系。
  • 在真实基因表达数据中,该方法成功识别出具有生物学合理性的信息流模式,包括通路间的时间延迟调控影响。
  • 该框架可为基因集定义网络中心性度量,基于总信息流识别关键的调控枢纽。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。