Skip to main content
QUICK REVIEW

[论文解读] Pathway-based feature selection algorithms identify genes discriminating patients with multiple sclerosis apart from controls

Lei Zhang, Linlin Wang|arXiv (Cornell University)|Aug 6, 2015
Gene expression and cancer classification参考文献 44被引用 4
一句话总结

本研究提出了一种改进的SAM-GSR特征选择算法,通过整合通路信息,利用微阵列数据识别多发性硬化症(MS)患者与健康对照组之间的差异基因。通过将生物通路作为先验知识,该方法在独立测试集上实现了高分类准确率,表明基于通路的特征选择可显著提升复杂基因组数据中的判别能力。

ABSTRACT

Introduction The focus of analyzing data from microarray experiments and extracting biological insight from such data has experienced a shift from identification of individual genes in association with a phenotype to that of biological pathways or gene sets. Meanwhile, feature selection algorithm becomes imperative to cope with the high dimensional nature of many modeling tasks in bioinformatics. Many feature selection algorithms use information contained within a gene set as a biological priori, and select relevant features by incorporating such information. Thus, an integration of gene set analysis with feature selection is highly desired. Significance analysis of microarray to gene-set reduction analysis (SAM-GSR) algorithm is a novel direction of gene set analysis, aiming at further reduction of gene set into a core subset. Here, we explore the feature selection trait possessed by SAM-GSR and then modify SAM-GSR specifically to better fulfill this role. Results and Conclusions Training on a multiple sclerosis (MS) microarray data using both SAM-GSR and our modification of SAM-GSR, excellent discriminative performance on an independent test set was achieved. To conclude, absorbing biological information from a gene set may be helpful for classification and feature selection. Discussion Given the fact the complete pathway information is far from completeness, a statistical method capable of constructing biologically meaningful gene networks is in demand. The basic requirement is that interplay among genes must be taken into account.

研究动机与目标

  • 为解决在高维基因组数据中识别MS相关基因的挑战。
  • 通过整合具有生物意义的基因集作为先验知识,改进特征选择方法。
  • 对SAM-GSR算法进行改进,以提升在MS研究中的分类性能。
  • 评估是否基于通路信息的特征选择可产生更稳健且具有生物学意义的基因特征。

提出的方法

  • 本研究对SAM-GSR算法进行修改,优先考虑特征选择而非基因集富集,重点提升分类性能。
  • 利用来自精心整理的基因集的通路信息作为生物学先验,指导特征选择过程。
  • 通过统计显著性检验,识别出在通路中能最好区分MS患者与对照组的核心基因子集。
  • 该算法在公开的MS微阵列数据集上进行训练,并在独立测试集上进行验证。
  • 通过整合显著性分析得到的p值与基因集隶属关系,执行特征选择,突出位于生物学相关通路中的基因。
  • 使用标准指标在保留的测试集上评估分类性能,以评估模型的泛化能力。

实验结果

研究问题

  • RQ1与传统方法相比,基于通路的特征选择是否能提升MS相关基因的识别能力?
  • RQ2在MS微阵列数据中,将生物通路知识整合到特征选择中是否能提升分类准确率?
  • RQ3改进后的SAM-GSR算法与原始算法相比,在判别性能上表现如何?
  • RQ4在通路中,哪些核心基因子集最能有效区分MS患者与对照组?

主要发现

  • 改进后的SAM-GSR算法在独立测试集上表现出优异的判别性能,表明其具有强大的泛化能力。
  • 整合通路信息显著提升了特征选择效果,使重点集中于具有生物学意义的基因。
  • 该方法成功地将基因集精简为核心子集,其中包含高度判别性的基因,从而增强了结果的可解释性。
  • 结果表明,诸如基因集等生物先验知识可显著提升高维基因组数据中的分类性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。