Skip to main content
QUICK REVIEW

[论文解读] Systematic Analysis of Cluster Similarity Indices: How to Validate Validation Measures

Martijn Gösgens, Alexey Tikhonov|arXiv (Cornell University)|Nov 12, 2019
Complex Network Analysis Techniques参考文献 25被引用 9
一句话总结

本文提出了一套严格的理论框架,用于评估和选择聚类相似性指数,通过定义和分析诸如恒定基线和单调性等理想属性。研究发现,只有皮尔逊积矩相关系数(Correlation Coefficient)和Sokal与Sneath提出的第一种指数满足所有关键属性,从而挑战了NMI和AMI等流行指标的广泛使用,这些指标在实践中被证明存在偏差且不一致。

ABSTRACT

Many cluster similarity indices are used to evaluate clustering algorithms, and choosing the best one for a particular task remains an open problem. We demonstrate that this problem is crucial: there are many disagreements among the indices, these disagreements do affect which algorithms are preferred in applications, and this can lead to degraded performance in real-world systems. We propose a theoretical framework to tackle this problem: we develop a list of desirable properties and conduct an extensive theoretical analysis to verify which indices satisfy them. This allows for making an informed choice: given a particular application, one can first select properties that are desirable for the task and then identify indices satisfying these. Our work unifies and considerably extends existing attempts at analyzing cluster similarity indices: we introduce new properties, formalize existing ones, and mathematically prove or disprove each property for an extensive list of validation indices. This broader and more rigorous approach leads to recommendations that considerably differ from how validation indices are currently being chosen by practitioners. Some of the most popular indices are even shown to be dominated by previously overlooked ones.

研究动机与目标

  • 为解决聚类相似性指数在性能排名上表现不一致的问题,该问题会降低实际系统性能。
  • 识别并形式化聚类相似性指数的理想理论属性,特别是恒定基线和单调性。
  • 提供一种基于理论的严谨方法,根据应用需求选择最合适的验证指标。
  • 通过揭示NMI和AMI等广泛使用指标的理论缺陷和偏差,挑战其主导地位。
  • 将先前的实证分析统一并扩展为对一系列指标(包括成对计数和相关性度量)的正式数学证明。

提出的方法

  • 形式化12项关键属性(包括恒定基线、单调性和对称性),作为聚类相似性指数的理论验证标准。
  • 开展全面的理论分析,数学上证明或证伪14种广泛使用的聚类相似性指数的每一项属性。
  • 提出成对计数指数的渐近恒定基线概念,实现对不同聚类规模下基线行为的严格评估。
  • 强化单调性的概念,统一并扩展先前的定义,确保在聚类细化过程中的稳定性。
  • 利用真实世界数据集和一个生产级新闻聚合系统,通过实证方法验证理论发现,并展示指标间的不一致性。
  • 将该框架应用于离线排序聚类算法,并与在线A/B测试结果对比,使理论结论与真实用户行为相契合。

实验结果

研究问题

  • RQ1哪些聚类相似性指数满足最理想的理论属性,如恒定基线和单调性?
  • RQ2在正式的理论审视下,NMI、AMI和Rand等流行指标相较于较少见的指标表现如何?
  • RQ3不同指标之间的分歧在多大程度上导致了实际应用中算法排名的不一致?
  • RQ4能否构建一个统一的理论框架,以根据应用需求指导验证指标的选择?
  • RQ5通过在线A/B实验验证,AMI和NMI max中的偏差在多大程度上影响了实际系统性能?

主要发现

  • 在14种聚类相似性指数中,仅有皮尔逊积矩相关系数(Correlation Coefficient)和Sokal与Sneath提出的第一种指数满足除“作为距离”之外的所有理想理论属性。
  • 皮尔逊积矩相关系数被证明具有渐近恒定基线特性,且可通过反余弦函数轻松转换为距离,因此在验证中极具适用性。
  • NMI max和NMI等流行指标始终偏好具有更多聚类数的划分,即使这些划分与人工标注的黄金标准对齐度更低。
  • 在一个真实世界的新闻聚合系统中,线上A/B测试证实,CC和S&S的排名与用户实际体验一致,而AMI错误地偏好了一个次优算法。
  • Rand指数在全部10个测试数据集中均偏好k=2个聚类,而NMI和NMI max在8–9个数据集中偏好k=2×ref-clusters,显示出系统性偏差。
  • AM I和F-measure等指标对小聚类敏感,这在聚类大小需保持平衡的应用中可能造成不利影响。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。