Skip to main content
QUICK REVIEW

[论文解读] A General Framework for Mixed Graphical Models

Eunho Yang, Pradeep Ravikumar|arXiv (Cornell University)|Nov 2, 2014
Soil Geostatistics and Mapping参考文献 49被引用 14
一句话总结

本文提出了一种新型多元分布类——块有向马尔可夫随机场(BDMRFs),通过结合有向与无向边的混合图结构,对二值、计数和连续等混合类型变量进行建模。该框架利用单变量指数族并借助块DAG进行串联,实现了高维混合数据中复杂依赖结构的可扩展、具有统计保证的估计,在基因组学及其他大数据应用中表现出色。

ABSTRACT

"Mixed Data" comprising a large number of heterogeneous variables (e.g. count, binary, continuous, skewed continuous, among other data types) are prevalent in varied areas such as genomics and proteomics, imaging genetics, national security, social networking, and Internet advertising. There have been limited efforts at statistically modeling such mixed data jointly, in part because of the lack of computationally amenable multivariate distributions that can capture direct dependencies between such mixed variables of different types. In this paper, we address this by introducing a novel class of Block Directed Markov Random Fields (BDMRFs). Using the basic building block of node-conditional univariate exponential families from Yang et al. (2012), we introduce a class of mixed conditional random field distributions, that are then chained according to a block-directed acyclic graph to form our class of Block Directed Markov Random Fields (BDMRFs). The Markov independence graph structure underlying a BDMRF thus has both directed and undirected edges. We introduce conditions under which these distributions exist and are normalizable, study several instances of our models, and propose scalable penalized conditional likelihood estimators with statistical guarantees for recovering the underlying network structure. Simulations as well as an application to learning mixed genomic networks from next generation sequencing expression data and mutation data demonstrate the versatility of our methods.

研究动机与目标

  • 为解决缺乏可计算的多元分布以建模异质数据类型(如二值、计数和连续变量)之间直接依赖关系的问题。
  • 开发一种通用的参数化多元分布类,支持在高维设置下对混合变量类型之间的丰富依赖结构进行建模。
  • 提供一种具有强统计保证的可扩展惩罚条件似然估计器,用于学习混合图模型中的潜在网络结构。
  • 实现对混合组学数据(如基因表达、突变和甲基化)的整合分析,以揭示具有生物学意义的调控关系。
  • 通过在统一框架中允许混合变量类型以及有向与无向依赖关系,推广现有图模型(如高斯模型和伊辛模型)

提出的方法

  • 提出块有向马尔可夫随机场(BDMRFs)作为一类基于节点条件单变量指数族构建的多元分布。
  • 使用块有向无环图(DAG)表示变量块之间的条件依赖关系,块间为有向边,块内为无向边。
  • 通过串联指数族的条件分布构建联合分布,在较宽松条件下确保可归一化性。
  • 采用惩罚条件似然估计器以恢复潜在网络结构,并提供一致性和稀疏性的理论保证。
  • 通过在单一集成模型中对混合基因组变量(如SNP,二值;RNA-seq,计数;甲基化,连续)进行建模,将该框架应用于真实世界数据。
  • 通过模拟实验和基于下一代测序的乳腺癌基因组数据真实应用验证该方法。

实验结果

研究问题

  • RQ1是否存在统一的统计框架,能够在单一多元分布中对混合类型变量(如二值、计数和连续)之间的复杂依赖关系进行建模?
  • RQ2在混合变量类型上基于条件指数族构建联合分布时,何种条件可确保其可归一化?
  • RQ3是否可实现一种可扩展的惩罚似然估计器,在高维混合数据中以理论保证恢复真实的潜在网络结构?
  • RQ4所提出的BDMRF框架在识别不同类型基因组生物标志物之间具有生物学意义的调控关系方面表现如何?
  • RQ5当变量块上的潜在DAG结构未知或部分误设时,该模型在多大程度上仍能学习到正确的依赖关系?

主要发现

  • BDMRF框架可在单一统一的概率模型中支持广泛的混合变量类型,包括二值、计数、连续及偏态连续变量。
  • 所提出的模型在相当宽松的条件下具备可归一化性,使其适用于多样化的统计应用。
  • 惩罚条件似然估计器在高维设置下以高概率一致地恢复潜在网络结构。
  • 在乳腺癌基因组学应用中,模型成功识别出已知的生物学关系,如TP53突变影响ADAM6表达,以及FGF3突变影响CCND1表达。
  • 在模拟实验中,该方法表现出色,能准确恢复混合图结构中的有向与无向边。
  • 该框架推广了现有模型(如高斯和伊辛图模型),将其扩展至支持混合数据类型及有向与无向依赖关系的统一框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。