Skip to main content
QUICK REVIEW

[论文解读] Comparison of Clustering Methods for Time Course Genomic Data: Applications to Aging Effects

Yafeng Zhang, Steve Horvath|arXiv (Cornell University)|Apr 29, 2014
Gene expression and cancer classification参考文献 27被引用 3
一句话总结

本研究比较了六种聚类方法——三种基于模型的方法(MFDA、FCM、MCLUST)和三种基于距离的方法(WGCNA、DTW、ACF)——在时间序列基因组数据上的表现,评估其在模拟数据和真实衰老相关微阵列数据集上的性能。WGCNA 和 MCLUST 表现最佳,其中 WGCNA 在生物相关性和准确性方面优于 MCLUST,尤其在长而密集的基因表达曲线中表现更优;而 MCLUST 提供基于模型的推断和不确定性估计。

ABSTRACT

Time course microarray data provide insight about dynamic biological processes. While several clustering methods have been proposed for the analysis of these data structures, comparison and selection of appropriate clustering methods are seldom discussed. We compared $3$ probabilistic based clustering methods and $3$ distance based clustering methods for time course microarray data. Among probabilistic methods, we considered: smoothing spline clustering also known as model based functional data analysis (MFDA), functional clustering models for sparsely sampled data (FCM) and model-based clustering (MCLUST). Among distance based methods, we considered: weighted gene co-expression network analysis (WGCNA), clustering with dynamic time warping distance (DTW) and clustering with autocorrelation based distance (ACF). We studied these algorithms in both simulated settings and case study data. Our investigations showed that FCM performed very well when gene curves were short and sparse. DTW and WGCNA performed well when gene curves were medium or long ($>=10$ observations). SSC performed very well when there were clusters of gene curves similar to one another. Overall, ACF performed poorly in these applications. In terms of computation time, FCM, SSC and DTW were considerably slower than MCLUST and WGCNA. WGCNA outperformed MCLUST by generating more accurate and biological meaningful clustering results. WGCNA and MCLUST are the best methods among the 6 methods compared, when performance and computation time are both taken into account. WGCNA outperforms MCLUST, but MCLUST provides model based inference and uncertainty measure of clustering results.

研究动机与目标

  • 评估并比较六种聚类方法在时间序列基因组数据上的表现,重点关注衰老相关动态特征。
  • 识别在不同曲线长度和采样模式下,功能性基因组数据中最准确且计算高效的聚类方法。
  • 评估真实衰老数据集中聚类结果的生物相关性,特别是识别与年龄相关的基因表达模式。
  • 考察基于模型的推断(如不确定性量化)与基于距离的效率之间在时间序列聚类中的权衡。

提出的方法

  • 评估三种基于模型的聚类方法:样条平滑聚类(MFDA)、稀疏采样数据的功能聚类(FCM)和基于模型的聚类(MCLUST),均使用混合模型和期望最大化(EM)算法进行参数估计。
  • 评估三种基于距离的方法:加权基因共表达网络分析(WGCNA)、动态时间规整(DTW)和基于自相关性的距离(ACF),每种方法使用不同的相似性度量来处理功能数据。
  • 应用样条平滑和分箱方法将横断面数据转换为伪纵向曲线,以确保各方法的一致性分析。
  • 在基于模型的方法中,使用贝叶斯信息准则(BIC)进行模型选择和聚类数量确定。
  • 通过模拟研究测试方法在不同曲线长度(短、中、长)和采样稀疏性下的鲁棒性。
  • 在真实数据集上验证结果:荷兰表达数据集和 SAFHS 表达数据集,重点关注衰老相关基因聚类。

实验结果

研究问题

  • RQ1哪种聚类方法在识别时间序列微阵列数据中具有生物意义的与年龄相关的基因表达模式方面表现最佳?
  • RQ2当基因曲线为短曲线、稀疏采样或长而密集采样时,聚类方法的表现如何比较?
  • RQ3在所评估的六种方法中,计算效率与聚类准确性的权衡如何?
  • RQ4基于模型的方法(如 MCLUST)与基于距离的方法(如 WGCNA)在提供不确定性估计和基于模型的推断方面有何差异?
  • RQ5不同的数据转换方法(平滑 vs. 分箱)在多大程度上影响聚类结果和方法性能?

主要发现

  • WGCNA 在识别具有清晰单调表达趋势(如随年龄下降)的衰老相关基因聚类方面优于所有其他方法,尤其在长而密集的曲线(≥15 个观测值)上表现更优。
  • FCM 在短而稀疏采样的曲线中表现最佳,这与其专为不规则采样功能数据设计的特性一致。
  • DTW 在中等长度曲线(10 个观测值)上表现强劲,尤其在曲线形状略有差异时。
  • ACF 在所有模拟和真实数据设置中表现均较差,准确率低且聚类分离效果差。
  • 在计算效率方面,WGCNA 和 MCLUST 显著快于 MFDA、FCM 和 DTW——在约 100 分钟内完成对 20,000 个基因的分析,而 FCM 和 MFDA 所需时间超过一天。
  • 尽管 WGCNA 识别的特征基因和平均曲线形状与其它方法相似,但其聚类在热图可视化中与年龄的关联性更强、更清晰,表明其具有更高的生物可解释性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。