Skip to main content
QUICK REVIEW

[论文解读] Nonlinear network-based quantitative trait prediction from transcriptomic data

Émilie Devijver, Mélina Gallopin|arXiv (Cornell University)|Jan 26, 2017
Bioinformatics and Genomic Networks参考文献 20被引用 6
一句话总结

该论文提出了一种完全参数化且可解释的模型——BLLiM,通过结合个体聚类与基因调控网络推断,从转录组数据预测定量性状。该模型在高斯混合回归中采用块对角协方差结构来建模非线性关系,在果蝇嗅觉行为数据上实现了具有竞争力的预测性能,并揭示了具有生物学意义的基因网络。

ABSTRACT

Quantitatively predicting phenotype variables by the expression changes in a set of candidate genes is of great interest in molecular biology but it is also a challenging task for several reasons. First, the collected biological observations might be heterogeneous and correspond to different biological mechanisms. Secondly, the gene expression variables used to predict the phenotype are potentially highly correlated since genes interact though unknown regulatory networks. In this paper, we present a novel approach designed to predict quantitative trait from transcriptomic data, taking into account the heterogeneity in biological samples and the hidden gene regulatory networks underlying different biological mechanisms. The proposed model performs well on prediction but it is also fully parametric, which facilitates the downstream biological interpretation. The model provides clusters of individuals based on the relation between gene expression data and the phenotype, and also leads to infer a gene regulatory network specific for each cluster of individuals. We perform numerical simulations to demonstrate that our model is competitive with other prediction models, and we demonstrate the predictive performance and the interpretability of our model to predict alcohol sensitivity from transcriptomic data on real data from Drosophila Melanogaster Genetic Reference Panel (DGRP).

研究动机与目标

  • 为解决从具有隐藏生物学异质性的高维相关转录组数据中预测连续定量性状的挑战。
  • 开发一种完全参数化的模型,实现准确预测与生物学可解释性,克服黑箱机器学习方法的局限性。
  • 推断与个体聚类相关的特定基因调控网络,捕捉表型变异背后的上下文特异性调控机制。
  • 在真实世界数据上展示该方法的实用性,特别是在理解果蝇(*Drosophila melanogaster*)自然变异中的嗅觉行为方面。
  • 提供一种可扩展至其他“组学”数据及具有潜在结构的异质生物学数据集的框架。

提出的方法

  • 使用带有隐变量的高斯混合线性回归模型,基于转录组-表型关系对个体进行聚类。
  • 采用逆回归(特别是 Sliced Inverse Regression 的扩展)来建模基因表达与定量性状之间的非线性关系。
  • 在条件协方差矩阵中施加块对角结构,以在每个个体聚类内编码基因调控网络。
  • 通过最大似然估计与期望最大化(EM)算法联合估计聚类成员、回归系数与网络结构。
  • 应用斜率启发式方法进行模型选择,以确定最优聚类数与网络复杂度。
  • 使用 Cytoscape 可视化跨聚类合并的基因网络,突出显示共享与聚类特异性相互作用。

实验结果

研究问题

  • RQ1完全参数化模型是否能有效捕捉转录组数据与定量性状之间的非线性关系,同时保持可解释性?
  • RQ2基于转录组-表型关联对个体进行聚类,相比全局模型,是否能显著提升预测准确性?
  • RQ3是否能可靠地推断出每个聚类内的基因调控网络?这些网络是否揭示了共调控基因的生物学上合理的模块?
  • RQ4在具有复杂表型变异的真实转录组数据上,该模型是否优于现有预测方法,包括机器学习技术?
  • RQ5推断出的网络与聚类在多大程度上能揭示果蝇嗅觉行为自然变异背后的分子机制?

主要发现

  • BLLiM 模型在模拟数据与真实数据上均实现了具有竞争力的预测性能,优于标准线性模型及其他非线性替代方法。
  • 该方法成功基于转录组-表型关系识别出果蝇个体的三个显著不同的聚类,每个聚类对应一个独特的基因调控网络。
  • 推断出的基因网络显示出生物学上合理的模块结构:部分基因与相互作用在聚类间共享(网络可视化中以黄色标示),而其他基因与相互作用则具有聚类特异性(以蓝色、绿色、红色标示)。
  • 估计的回归系数矩阵呈现出稀疏结构,表明每个聚类中仅有少数基因具有预测能力,从而增强了模型的可解释性。
  • 该模型揭示果蝇嗅觉行为的变异与不同的调控程序相关,关键基因如 *foraging* 和 *takeout* 在聚类特异性网络中被识别。
  • 该框架具有可扩展性,可应用于其他组学数据类型,包括代谢组学与蛋白质组学数据,并可扩展至多组学整合。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。