Skip to main content
QUICK REVIEW

[论文解读] TREEOME: A framework for epigenetic and transcriptomic data integration to explore regulatory interactions controlling transcription

David Budden, Daniel Hurley|arXiv (Cornell University)|Feb 5, 2015
Epigenetics and DNA Methylation参考文献 22被引用 7
一句话总结

TREEOME 是一种基于决策树的框架,通过整合基因水平的 DNA 甲基化和组蛋白修饰数据,将基因划分为调控类别,以建模条件性和协同性表观遗传相互作用,从而提高全基因组转录本丰度的预测能力。当使用每位置平均甲基化分数(MMFS)作为甲基化度量指标时,其预测值与实测值之间的转录本丰度皮尔逊相关系数 r > 0.70,调整决定系数 adj.R² 提高了 0.06。

ABSTRACT

Motivation: Predictive modelling of gene expression is a powerful framework for the in silico exploration of transcriptional regulatory interactions through the integration of high-throughput -omics data. A major limitation of previous approaches is their inability to handle conditional and synergistic interactions that emerge when collectively analysing genes subject to different regulatory mechanisms. This limitation reduces overall predictive power and thus the reliability of downstream biological inference. Results: We introduce an analytical modelling framework (TREEOME: tree of models of expression) that integrates epigenetic and transcriptomic data by separating genes into putative regulatory classes. Current predictive modelling approaches have found both DNA methylation and histone modification epigenetic data to provide little or no improvement in accuracy of prediction of transcript abundance despite, for example, distinct anti-correlation between mRNA levels and promoter-localised DNA methylation. To improve on this, in TREEOME we evaluate four possible methods of formulating gene-level DNA methylation metrics, which provide a foundation for identifying gene-level methylation events and subsequent differential analysis, whereas most previous techniques operate at the level of individual CpG dinucleotides. We demonstrate TREEOME by integrating gene-level DNA methylation (bisulfite-seq) and histone modification (ChIP-seq) data to accurately predict genome-wide mRNA transcript abundance (RNA-seq) for H1-hESC and GM12878 cell lines. Availability: TREEOME is implemented using open-source software and made available as a pre-configured bootable reference environment. All scripts and data presented in this study are available online at http://sourceforge.net/projects/budden2015treeome/.

研究动机与目标

  • 为克服传统预测模型在处理表观遗传标记之间条件性和协同性调控相互作用方面的局限性。
  • 开发一种通过整合基因水平的 DNA 甲基化和组蛋白修饰数据来提升转录本丰度预测能力的框架。
  • 评估并识别用于预测建模的最有效的基因水平 DNA 甲基化度量指标。
  • 展示 TREEOME 在 H1-hESC 和 GM12878 细胞系中准确预测 RNA-seq 表达水平的实用性。
  • 基于表观遗传特征,实现生物上有意义的、无监督的基因调控类别分类。

提出的方法

  • 使用基于决策树的分析框架(TREEOME),根据表观遗传特征将基因分类为潜在的调控类别。
  • 采用四种基因水平的 DNA 甲基化度量指标——MMFS、MMS、MPM 和 MPMF,以量化启动子区域的甲基化水平,并评估其预测能力。
  • 整合 H3K4me3、H3K27me3、H3K9me3 和 H2A.Z 的 ChIP-seq 数据,以建模上下文敏感的调控相互作用。
  • 应用无监督的阈值选择方法来定义调控类别,优先保证生物可解释性,而非纯粹的统计优化。
  • 在每个调控类别内使用线性回归模型,基于整合的表观遗传特征预测转录本丰度。
  • 使用预测值与实测 RNA-seq 表达之间的皮尔逊相关系数(r)和调整决定系数(adj.R²)验证模型性能。

实验结果

研究问题

  • RQ1哪种基因水平的 DNA 甲基化度量指标与转录本丰度相关性最强,并能提升预测模型的准确性?
  • RQ2基于表观遗传特征将基因划分为调控类别,是否能提升转录本丰度的预测能力?
  • RQ3组蛋白修饰之间的条件性相互作用(如 H3K4me3、H3K27me3、H2A.Z)如何影响转录输出?
  • RQ4在存在复杂染色质调控的情况下,整合 DNA 甲基化数据在多大程度上能提高转录本丰度预测的准确性?
  • RQ5TREEOME 是否能够解析具有 H2A.Z 的基因中的调控异质性,这些基因表现出上下文依赖的转录效应?

主要发现

  • 在 H1-hESC 和 GM12878 细胞系中,每位置平均甲基化分数(MMFS)度量指标与基因表达水平表现出最强的负相关性。
  • 当使用 MMFS 时,TREEOME 的预测准确度提升了 Δadj.R² = 0.06,预测值与实测值之间的转录本丰度皮尔逊相关系数 r > 0.70。
  • H2A.Z− 基因的预测准确度显著更高,而 H2A.Z+ 基因的准确度较低,可能归因于其调控异质性。
  • 该框架首次成功实现了全基因组范围内亚硫酸盐测序(DNA 甲基化)、ChIP-seq(组蛋白修饰)和 RNA-seq(转录本丰度)数据的整合预测。
  • 将基因划分为调控类别,使得能够建模条件性和协同性相互作用,例如在 H2A.Z 和 H3K27me3 存在或缺失的情况下,H3K4me3 的上下文依赖性作用。
  • TREEOME 允许为甲基化与未甲基化基因推导出不同的调控角色及其置信度值,从而在预测工作流程中提升生物可解释性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。