Skip to main content
QUICK REVIEW

[论文解读] Identification of essential and functionally moduled genes through the microarray assay

Kyoohyoung Rho, Hye Gwang Jeong|arXiv (Cornell University)|Jan 9, 2003
Bioinformatics and Genomic Networks参考文献 2被引用 3
一句话总结

该论文提出了一种无阈值、自组织的体外方法,利用微阵列数据识别关键基因和功能模块化基因。通过构建基于皮尔逊相关系数的基因转录网络,并分析临界分数 $p_m$(最大簇)和 $p_s$(幂律连通性)处的渗滤转变,该方法识别出在 $p_s$ 处平均连通性最高的簇,该簇最有可能包含关键基因——在 *S. cerevisiae* 中,64个基因的簇中关键性达到73%,且具有强烈的功能一致性。

ABSTRACT

Identification of essential genes is one of the ultimate goals of drug designs. Here we introduce an {\it in silico} method to select essential genes through the microarray assay. We construct a graph of genes, called the gene transcription network, based on the Pearson correlation coefficient of the microarray expression level. Links are connected between genes following the order of the pair-wise correlation coefficients. We find that there exist two meaningful fractions of links connected, $p_m$ and $p_s$, where the number of clusters becomes maximum and the connectivity distribution follows a power law, respectively. Interestingly, one of clusters at $p_m$ contains a high density of essential genes having almost the same functionality. Thus the deletion of all genes belonging to that cluster can lead to lethal inviable mutant efficiently. Such an essential cluster can be identified in a self-organized way. Once we measure the connectivity of each gene at $p_s$. Then using the property that the essential genes are likely to have more connectivity, we can identify the essential cluster by finding the one having the largest mean connectivity per gene at $p_m$.

研究动机与目标

  • 克服传统依赖阈值的聚类方法在从微阵列数据中识别关键基因时的局限性。
  • 不单独识别关键基因,而是将其作为功能一致簇的一部分,以提高生物相关性和预测准确性。
  • 开发一种无需调节参数的自组织方法,避免在簇选择中引入人为偏差。
  • 将转录网络中的基因连通性与关键性联系起来,利用高度连通的基因更可能为关键基因的原理。
  • 基于其簇的主要功能模块,为未注释基因分配推测功能。

提出的方法

  • 通过按微阵列表达水平皮尔逊相关系数降序排列基因对,构建基因转录网络。
  • 识别 $p_m$,即簇数(含≥2个基因)最大的分数,表示功能模块性。
  • 识别 $p_s$,即连通性分布符合幂律的分数,表明尺度自由网络结构的出现。
  • 在 $p_s$ 处测量每个基因的连通性 $k_i(p_s)$,以评估其在网络中的中心性。
  • 返回至 $p_m$ 处的网络,计算每个簇 $J$ 的平均连通性 $\langle k^J \rangle = \frac{\sum_{i \in J} k_i(p_s)}{N^J(p_m)}$。
  • 基于假设关键基因高度连通,选择 $\langle k^J \rangle$ 最高的簇作为最可能包含关键基因的簇。

实验结果

研究问题

  • RQ1通过聚焦于功能簇而非单个基因,能否更有效地识别关键基因?
  • RQ2无用户定义阈值的自组织网络方法是否能改善微阵列数据中关键基因的识别?
  • RQ3转录网络中的高连通性与基因关键性之间是否存在相关性?
  • RQ4在 $p_m$ 和 $p_s$ 处识别出的簇是否表现出功能一致性,表明模块化组织?
  • RQ5能否基于簇的主要功能类别可靠地为未注释基因分配推测功能?

主要发现

  • 在 $p_s$ 处平均连通性最高的簇中,64个基因中有47个为关键基因,关键性达到73%。
  • 在 $p_m$ 处第三大簇被识别为关键簇,包含64个基因,其中47个为已知关键基因。
  • $p_m$ 处的功能聚类显示出强烈的同质性:最大簇富含氨基酸代谢,第二簇富含小分子转运,第三簇富含RNA加工,第四簇富含蛋白质合成。
  • $\langle k^J \rangle$ 与关键性分数 $\mathcal{E}^J$ 之间的相关性强烈且单调,验证了该方法的预测能力。
  • 该方法成功基于其簇的主要功能类别,为未注释基因分配了推测功能,如表1所示。
  • 该方法避免了任意阈值,为从微阵列数据中识别关键、功能模块化基因提供了一种自组织、可重复的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。