Skip to main content
QUICK REVIEW

[论文解读] Clustering functional data using wavelets

Anestis Antoniadis, Xavier Brossat|Jan 25, 2011
Time Series Analysis and Forecasting被引用 12
一句话总结

本文提出了两种基于小波的聚类方法,用于高维功能时间序列聚类,利用多分辨率分析检测局部模式并分组相似曲线。第一种方法使用小波能量特征结合k-means聚类;第二种方法采用基于小波相干性的差异性度量并结合k中心点聚类。两种方法在法国电力需求数据上的表现均优于传统的L2基聚类方法,揭示了细微的、基于形状的聚类,包括罕见的假日模式。

ABSTRACT

We present two methods for detecting patterns and clusters in high dimensional time-dependent functional data. Our methods are based on wavelet-based similarity measures, since wavelets are well suited for identifying highly discriminant local time and scale features. The multiresolution aspect of the wavelet transform provides a time-scale decomposition of the signals allowing to visualize and to cluster the functional data into homogeneous groups. For each input function, through its empirical orthogonal wavelet transform the first method uses the distribution of energy across scales generate a handy number of features that can be sufficient to still make the signals well distinguishable. Our new similarity measure combined with an efficient feature selection technique in the wavelet domain is then used within more or less classical clustering algorithms to effectively differentiate among high dimensional populations. The second method uses dissimilarity measures between the whole time-scale representations and are based on wavelet-coherence tools. The clustering is then performed using a k-centroid algorithm starting from these dissimilarities. Practical performance of these methods that jointly designs both the feature selection in the wavelet domain and the classification distance is demonstrated through simulations as well as daily profiles of the French electricity power demand.

研究动机与目标

  • 为解决非平稳功能时间序列聚类的挑战,特别是在电力需求预测中,标准方法难以捕捉基于形状的模式。
  • 开发一种聚类框架,以保留对识别功能数据中不同行为模式至关重要的局部时频特征。
  • 改进传统基于L2的聚类方法,后者通常仅根据均值水平分组曲线,而忽略形状差异。
  • 通过在小波域内集成特征选择,实现可解释性更强、更具适应性的聚类,以提升模型性能与洞察力。

提出的方法

  • 第一种方法应用离散小波变换(DWT)提取各尺度上的能量分布,降低维度的同时保留具有区分性的局部特征。
  • 在小波域中采用特征选择技术,仅保留具有信息量的系数,随后应用经典的k-means聚类算法。
  • 第二种方法计算曲线完整时频表示之间的基于小波相干性的差异度量,捕捉相位与振幅关系。
  • 利用这些差异度量,应用k中心点算法,基于小波域中的结构相似性形成聚类。
  • 小波变换实现多分辨率分解,使局部特征在时间和尺度上均可被检测。
  • 两种方法均在真实日电力需求曲线上进行验证,结果通过调整兰德指数与基于日历的标签进行比较。

实验结果

研究问题

  • RQ1与标准L2基向量聚类相比,基于小波的特征提取是否能提升高维功能数据的聚类性能?
  • RQ2基于小波相干性的差异度量在捕捉功能时间序列中的基于形状的相似性方面效果如何?
  • RQ3所提出的方法能否检测到罕见或异常的曲线模式(如特定假日或受天气影响的日子),而这些是传统聚类方法所遗漏的?
  • RQ4在小波域中进行特征选择在多大程度上提升了聚类的可解释性与准确性?
  • RQ5与经典聚类方法相比,基于小波的聚类结果在调整兰德指数和实际可解释性方面表现如何?

主要发现

  • 基于小波特征提取的方法与AC聚类方法相比,调整兰德指数为0.26,表明其与标准方法存在显著差异。
  • 基于小波相干性差异度量的方法与AC相比获得更高的调整兰德指数0.32,表明其与有意义分组的对齐性更好。
  • WER与MCA聚类结果之间的调整兰德指数为0.57,表明两种所提方法之间具有显著一致性。
  • 基于WER的方法成功识别出两个不同的‘特殊日子’聚类:一个仅包含圣诞节和元旦(2天),另一个包含冬季周六和夏令时调整日。
  • 基于MCA的方法检测到更广泛的特殊日子,包括除夕、平安夜、寒冷的周末,其中聚类G包含19个观测值。
  • 研究表明,基于小波的聚类优于传统的L2基聚类,因其能捕捉基于形状的模式,而不仅限于均值水平的差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。