[论文解读] Unsupervised learning of transcriptional regulatory networks via latent tree graphical models
本文提出一种潜在树图模型,用于从基因表达数据推断转录调控网络,而无需依赖转录因子(TF)mRNA水平作为调节活性的代理指标。通过将隐藏调节因子建模为潜变量并采用高效的无监督学习方法,该方法识别出共调控的基因群组,恢复了酿酒酵母应激反应中的已知调控关系,并预测了新型条件特异性TF活性,例如在渗透压应激下Msn4的结合活性。
Gene expression is a readily-observed quantification of transcriptional activity and cellular state that enables the recovery of the relationships between regulators and their target genes. Reconstructing transcriptional regulatory networks from gene expression data is a problem that has attracted much attention, but previous work often makes the simplifying (but unrealistic) assumption that regulator activity is represented by mRNA levels. We use a latent tree graphical model to analyze gene expression without relying on transcription factor expression as a proxy for regulator activity. The latent tree model is a type of Markov random field that includes both observed gene variables and latent (hidden) variables, which factorize on a Markov tree. Through efficient unsupervised learning approaches, we determine which groups of genes are co-regulated by hidden regulators and the activity levels of those regulators. Post-processing annotates many of these discovered latent variables as specific transcription factors or groups of transcription factors. Other latent variables do not necessarily represent physical regulators but instead reveal hidden structure in the gene expression such as shared biological function. We apply the latent tree graphical model to a yeast stress response dataset. In addition to novel predictions, such as condition-specific binding of the transcription factor Msn4, our model recovers many known aspects of the yeast regulatory network. These include groups of co-regulated genes, condition-specific regulator activity, and combinatorial regulation among transcription factors. The latent tree graphical model is a general approach for analyzing gene expression data that requires no prior knowledge of which possible regulators exist, regulator activity, or where transcription factors physically bind.
研究动机与目标
- 为解决现有方法假设TF mRNA水平反映TF活性的局限性,该假设在生物学上不准确,因存在转录后和翻译后调控。
- 开发一种无监督方法,从基因共表达数据中推断隐藏的转录调节因子,而无需事先了解调节因子、其活性或结合位点。
- 识别共调控的基因模块,并通过功能注释和基序富集分析为潜变量赋予生物学意义。
- 在酿酒酵母应激反应数据集上验证该方法,恢复已知的调控关系并预测新的调控关系,如条件特异性TF结合。
提出的方法
- 采用潜在树图模型,其中观测到的基因表达水平条件依赖于隐藏(潜伏)调节因子变量,并在树结构上进行因子分解。
- 应用Choi等人[23]提出的保证学习算法,从未知潜节点数量和位置的基因表达数据中恢复潜在的树结构和参数。
- 基于基因表达相关性距离,采用动态邻域选择方法定义每个潜节点的影响范围,使用可调参数λ = 0.15。
- 通过统计检验(Fisher精确检验和fdrtool进行贝叶斯FDR估计)评估潜节点基因集与已知TF结合位点之间的重叠情况。
- 使用基因本体论(GO)生物过程术语对潜节点进行注释,并利用WebMOTIFS进行从头基序发现,以将潜变量与已知TF结合基序关联。
- 利用渗透压应激样本中潜变量的条件均值对调节因子进行排序和识别,从而定义特定于渗透压应激的TF。
实验结果
研究问题
- RQ1潜在树图模型能否在不假设TF mRNA水平反映TF活性的前提下,从未表达数据中恢复出具有生物学意义的转录调节因子?
- RQ2该方法在酿酒酵母应激反应等已充分表征的生物系统中,识别已知调控关系和共调控基因模块的能力如何?
- RQ3通过基序富集分析和GO术语分析,潜变量在多大程度上可被注释为特定转录因子或功能基因群组?
- RQ4该方法能否预测通过标准mRNA方法无法检测到的新颖、条件特异性TF活性?
- RQ5与ARACNE等成熟方法相比,该潜在树模型在可扩展性和生物学相关性方面的表现如何?
主要发现
- 该方法成功恢复了酿酒酵母转录调控网络的已知特征,包括共调控的基因模块以及转录因子之间的组合性调控。
- 它预测了新颖的、条件特异性的TF活性,例如在渗透压应激下Msn4的结合活性,该活性仅通过mRNA表达无法检测到。
- 在山梨醇诱导的渗透压应激样本中,按条件均值排序的前50%潜节点显著与已知的渗透压应激TF相关,且该结果在不同阈值和样本子集下均保持稳健。
- 潜节点影响范围的扩展邻域与已知TF结合位点之间存在显著重叠,且通过贝叶斯FDR估计控制了统计显著性。
- 从头基序发现识别出了与潜节点相关基因集中序列模式匹配的已知酿酒酵母TF结合基序,支持了其生物学相关性。
- 潜在树模型在可扩展性方面优于ARACNE,在不到8小时内完成对1035个基因子集的分析,而ARACNE在合理时间内未能收敛于完整数据集。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。