[论文解读] Categorization of interestingness measures for knowledge extraction
本文提出了一种基于21个预定义理论属性的正式分类方法,利用两种聚类技术——Ward层次聚类和k-means——对61种关联规则有趣性度量进行分类。研究识别出七个不同的度量类别,为知识提取中的度量选择提供了系统性框架,通过将具有相似行为特性的度量归类,降低用户选择复杂度,提升规则相关性过滤效果。
Finding interesting association rules is an important and active research field in data mining. The algorithms of the Apriori family are based on two rule extraction measures, support and confidence. Although these two measures have the virtue of being algorithmically fast, they generate a prohibitive number of rules most of which are redundant and irrelevant. It is therefore necessary to use further measures which filter uninteresting rules. Many synthesis studies were then realized on the interestingness measures according to several points of view. Different reported studies have been carried out to identify "good" properties of rule extraction measures and these properties have been assessed on 61 measures. The purpose of this paper is twofold. First to extend the number of the measures and properties to be studied, in addition to the formalization of the properties proposed in the literature. Second, in the light of this formal study, to categorize the studied measures. This paper leads then to identify categories of measures in order to help the users to efficiently select an appropriate measure by choosing one or more measure(s) during the knowledge extraction process. The properties evaluation on the 61 measures has enabled us to identify 7 classes of measures, classes that we obtained using two different clustering techniques.
研究动机与目标
- 扩展并正式定义用于表征‘良好’有趣性度量的属性集合,适用于关联规则挖掘。
- 分析61种现有有趣性度量在这些形式化属性上的行为表现。
- 利用聚类技术识别具有相似属性特征的度量的自然分组(类别)。
- 为检测到的度量类别提供语义解释,以改善用户在选择相关度量时的指导。
- 将分类结果与先前研究进行验证,并提供基于共识的分类体系,以供数据挖掘中的实际应用。
提出的方法
- 正式定义21个被认为对有趣性度量而言理想的理论属性,例如对称性、单调性以及在某些变换下的不变性。
- 构建一个61种度量 × 21个属性的矩阵,其中每个条目表示某一特定度量是否满足给定属性。
- 应用Ward的层次聚类方法对矩阵进行聚类,基于属性相似性检测度量的嵌套分组。
- 采用改进的k-means聚类算法作为非层次聚类方法,以验证层次聚类结果。
- 通过两种聚类技术的一致性结果生成共识分类,以确保结果的稳健性。
- 对最终分类结果进行语义解释,将度量类别与其底层的统计或逻辑行为相关联。
实验结果
研究问题
- RQ1在61种有趣性度量中,有多少种在21个形式化属性上表现出一致的行为?
- RQ2聚类技术能否基于共享的属性特征揭示度量的自然分组?
- RQ3所生成的度量类别是否对应于规则挖掘中有意义的语义类别?
- RQ4不同算法(Ward与k-means)的聚类结果如何比较,是否存在共识?
- RQ5所识别的类别能否与现有文献进行验证,并用于实际中指导度量选择?
主要发现
- 通过共识聚类,本研究识别出七个不同的有趣性度量类别,每一类代表一组具有相似行为特性的度量。
- Ward方法与k-means的聚类结果表现出高度一致性,验证了所检测类别的稳健性。
- Lift、Kulczynski和Jaccard等度量被归入一个以对称性、基于置信度评估规则强度为特征的类别。
- 另一类别包含Piatetsky-Shapiro和Leverage等度量,其特点在于强调对独立性的偏离,对罕见事件敏感。
- 包含互信息(Mutual Information)和J-Measure的类别因其信息论基础以及对联合分布模式的敏感性而具有独特性。
- 该分类提供了一个语义框架,帮助用户根据其期望的规则过滤标准(如新颖性、强度或可靠性)选择合适的度量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。