[论文解读] Outcome signature genes in breast cancer: is there a unique set?
本研究调查了乳腺癌结果特征基因是否在不同分析中形成一组独特且稳定的基因集合。通过使用单一数据集和方法,研究显示所选基因列表对患者子集的选择极为敏感,揭示了由于强相关性波动和生存相关基因相关性微小差异,同一数据中可能涌现出多个预测性能相当的基因集合。
Motivation: Predicting the metastatic potential of primary malignant tissues has direct bearing on choice of therapy. Several microarray studies yielded gene sets whose expression profiles successfully predicted survival (Ramaswamy et al 2003; Sorlie et al 2001; van't Veer et al 2003). Nevertheless, the overlap between these gene sets is almost zero. Such small overlaps were observed also in other complex diseases (Lossos et al 2003; Miklos and Maleszka 2004), and the variables that could account for the differences had evoked a wide interest. One of the main open questions in this context is whether the disparity can be attributed only to trivial reasons such as different technologies, different patients and different types of analysis. Results: To answer this question we concentrated on one single breast cancer dataset, and analyzed it by one single method, the one which was used by van't Veer et al to produce a set of outcome predictive genes. We showed that in fact the resulting set of genes is not unique; it is strongly influenced by the subset of patients used for gene selection. Many equally predictive lists could have been produced from the same analysis. Three main properties of the data explain this sensitivity: (a) many genes are correlated with survival; (b) the differences between these correlations are small; (c) the correlations fluctuate strongly when measured over different subsets of patients. A possible biological explanation for these properties is discussed.
研究动机与目标
- 确定乳腺癌结果特征基因在生物学上是否真正独特,或仅是方法学上的不稳定性所致。
- 调查不同研究中基因集合的差异是否源于技术或患者选择等微小因素。
- 评估在单一数据集中使用一致方法时基因选择的稳定性。
- 识别导致预测基因列表非唯一性的数据特性。
提出的方法
- 将van't Veer等人所用的相同统计方法应用于单一乳腺癌芯片数据集。
- 系统性地改变用于基因选择的患者子集,以评估所得基因列表的敏感性。
- 测量在不同患者子集中,单个基因表达水平与患者生存之间的相关性强度。
- 分析这些相关性在随机与非随机患者子集之间的波动情况。
- 评估从不同子集导出的替代基因列表的预测性能,以检验其性能是否等价。
- 探讨观察到的基因选择不稳定性可能存在的生物学解释。
实验结果
研究问题
- RQ1乳腺癌的结果特征基因集合是否真正唯一?是否存在多个不同的基因列表可实现等效的预测性能?
- RQ2在相同数据集中,患者子集的选择在多大程度上影响了生存预测的基因选择?
- RQ3哪些数据特性导致了尽管分析方法一致,基因选择仍存在高度可变性?
- RQ4观察到的已发表基因集合之间缺乏重叠,是否可归因于统计不稳定性而非生物学差异?
- RQ5基因表达与生存之间的相关性是否足够稳健,以产生稳定且可重复的基因特征?
主要发现
- 相同的数据集和分析方法产生了多个不同的基因列表,且这些列表对患者生存的预测能力均等效。
- 结果特征基因的选择对分析中所用患者子集的具体构成极为敏感。
- 许多基因与生存表现出强烈但波动的相关性,其在不同子集中的相关性值存在微小差异。
- 基因间相关性差异的幅度虽小,但足以改变最终选定的基因列表。
- 基因选择的不稳定性主要源于在不同患者子集上计算相关性估计值时的高方差。
- 本研究表明,已发表基因集合之间缺乏重叠,可能源于方法学上的敏感性,而非生物学上的差异。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。