Skip to main content
QUICK REVIEW

[论文解读] How Many Genes Are Needed for a Discriminant Microarray Data Analysis ?

Wentian Li, Yaning Yang|ArXiv.org|Apr 6, 2001
Gene expression and cancer classification被引用 10
一句话总结

本研究利用白血病数据集(38个样本,7129个基因)调查了在微阵列数据中进行有效判别分析所需的最少基因数。通过AIC和BIC进行模型选择,发现仅需1–2个基因(特别是基因#4847,zyxin)即为最优,而使用相等或基于似然的权重进行模型平均时,超过10–25个基因后收益递减,表明在此数据集中先前研究使用的50个基因过多,原因在于存在强个体预测因子且样本量较小。

ABSTRACT

The analysis of the leukemia data from Whitehead/MIT group is a discriminant analysis (also called a supervised learning). Among thousands of genes whose expression levels are measured, not all are needed for discriminant analysis: a gene may either not contribute to the separation of two types of tissues/cancers, or it may be redundant because it is highly correlated with other genes. There are two theoretical frameworks in which variable selection (or gene selection in our case) can be addressed. The first is model selection, and the second is model averaging. We have carried out model selection using Akaike information criterion and Bayesian information criterion with logistic regression (discrimination, prediction, or classification) to determine the number of genes that provide the best model. These model selection criteria set upper limits of 22-25 and 12-13 genes for this data set with 38 samples, and the best model consists of only one (no.4847, zyxin) or two genes. We have also carried out model averaging over the best single-gene logistic predictors using three different weights: maximized likelihood, prediction rate on training set, and equal weight. We have observed that the performance of most of these weighted predictors on the testing set is gradually reduced as more genes are included, but a clear cutoff that separates good and bad prediction performance is not found.

研究动机与目标

  • 确定在样本量有限的微阵列数据中,进行准确判别分析所需的最少基因数。
  • 评估先前白血病分类研究中使用的50个基因是否必要或过多。
  • 比较模型选择(AIC/BIC)与模型平均方法在分类基因子集选择中的表现。
  • 评估基因冗余与强个体预测因子对模型复杂度与性能的影响。

提出的方法

  • 通过AIC和BIC进行模型选择,以识别逻辑回归模型中最佳基因数量。
  • 采用逐步变量选择法,评估基因子集的性能。
  • 使用三种加权方案进行模型平均:最大似然权重、训练集预测率权重和相等权重。
  • 在训练集和独立测试集上评估预测性能。
  • 通过权重计算有效基因数,以反映各模型的贡献程度。
  • 使用逻辑回归进行分类:P(AML) = 1 / (1 + exp(−a₀ − Σaⱼxⱼ)),其中基因表达水平xⱼ作为预测变量。

实验结果

研究问题

  • RQ1在仅含38个样本的Whitehead/MIT白血病微阵列数据集中,判别分析所需的最优基因数是多少?
  • RQ2在该小样本设置下,先前分类研究中使用50个基因是否属于过拟合,或属于必要的稳健性体现?
  • RQ3AIC与BIC模型选择准则在识别最佳基因子集用于分类方面表现如何比较?
  • RQ4使用多个基因进行模型平均是否能提升预测性能,超越单基因模型?
  • RQ5在模型平均中,当包含超过少数几个基因时,是否存在明确的性能截止点?

主要发现

  • 使用AIC和BIC进行模型选择表明,最优模型仅需1–2个基因,特别是基因#4847(zyxin),其在训练集上实现了完美分类。
  • 最佳单基因模型(基因#4847)在训练集上的预测率为36/38,在测试集上也为36/38,表明其具有良好的泛化能力。
  • 在模型平均中包含超过25个基因时,测试集上的性能下降,虽无明确截止点,但预测准确率呈持续下降趋势。
  • 模型平均中最大似然权重方案因主导基因的强烈影响,实质上退化为单基因模型,导致前10个基因的有效基因数仅为1.05。
  • 相等权重与基于预测率的模型平均方案表现相似,但当包含超过10–20个基因时,两者在测试集上的性能均出现下降。
  • 研究结论认为,对于此数据集,50个基因过多,因为强个体预测因子与小样本量使得更大的基因集不仅不必要,反而可能因过拟合而带来损害。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。