Skip to main content
QUICK REVIEW

[论文解读] Globally Homogenous Mixture Components and Local Heterogeneity of Rank Data

Vartan Choulakian|arXiv (Cornell University)|Aug 17, 2016
Consumer Market Behavior and Pricing参考文献 2被引用 4
一句话总结

本文提出一种基于方向性与逻辑的分析方法,利用带反向编码的出租车对应分析(TCA)检测排名数据中的全局同质混合成分与局部异质性。该方法引入全局同质性系数(GHC),量化块内排序一致性程度,在两个真实数据集中分别实现75.32%与80.36%的同质性水平,优于传统距离模型与潜在类别模型在检测小型、隐蔽群体方面的能力。

ABSTRACT

The traditional methods of finding mixture components of rank data are mostly based on distance and latent class models; these models may exhibit the phenomenon of masking of groups of small sizes; probably due to the spherical nature of rank data. Our approach diverges from the traditional methods; it is directional and uses a logical principle, the law of contradiction. We discuss the concept of a mixture for rank data essentially in terms of the notion of global homogeneity of its group components. Local heterogeneities may appear once the group components of the mixture have been discovered. This is done via the exploratory analysis of rank data by taxicab correspondence analysis with the nega coding: If the first factor is an affine function of the Borda count, then we say that the rank data are globally homogenous, and local heterogeneities may appear on the consequent factors; otherwise, the rank data either are globally homogenous with outliers, or a mixture of globally homogenous groups. Also we introduce a new coefficient of global homogeneity, GHC. GHC is based on the first taxicab dispersion measure: it takes values between 0 and 100\%, so it is easily interpretable. GHC measures the extent of crossing of scores of voters between two or three blocks seriation of the items where the Borda count statistic provides consensus ordering of the items on the first axis. Examples are provided. Key words: Preferences; rankings; Borda count; global homogeneity coefficient; nega coding; law of contradiction; mixture; outliers; taxicab correspondence analysis; masking.

研究动机与目标

  • 为解决传统混合模型在分析排名数据时因排名的球面对称性而难以检测小型群体的问题。
  • 基于矛盾律发展一种方向性方法,以检测排名数据中的全局同质成分。
  • 引入一种新的、可解释的系数GHC,基于第一轴离散度与块间得分交叉情况,量化全局同质性。
  • 在识别出全局同质群体后,通过高维TCA输出实现对局部异质性的探索性检测。
  • 展示该方法在识别真实数据集中小型、独特选民群体方面,优于距离模型与潜在类别模型。

提出的方法

  • 对反向编码的排名数据执行出租车对应分析(TCA),其中每项排名通过反向Borda得分转换,以增强方向敏感性。
  • 应用Borda计数法,确定沿第一TCA轴的项目共识排序,作为分块划分的参考基准。
  • 将全局同质性定义为第一TCA因子是Borda计数的仿射函数,表明块间无得分交叉。
  • 在反向编码数据上计算首个出租车离散度度量,以推导出全局同质性系数(GHC),其取值范围为0%至100%。
  • 运用矛盾律识别异常值,不将其视为数据点,而是识别出其偏好违背全局同质群体逻辑一致性的选民。
  • 基于第一TCA轴对项目进行分块排序,以评估块间得分交叉的程度,GHC即量化该现象。

实验结果

研究问题

  • RQ1基于方向性与逻辑的方法是否能比传统距离模型或潜在类别模型更有效地检测排名数据中的全局同质成分?
  • RQ2所提出的GHC系数在存在小型、隐蔽群体的情况下,能否准确衡量排名数据的全局同质性?
  • RQ3在识别出全局同质群体后,局部异质性如何在高维TCA输出中显现?
  • RQ4矛盾律是否可用于识别排名数据中的异常值,而无需依赖距离或模型假设?
  • RQ5与标准方法相比,TCA中的反向编码在检测球面对称排名数据的结构层次方面有何改进?

主要发现

  • 在Croon的政治目标数据集中,该方法识别出16名心理学家构成的全局同质群体,GHC为75.32%,揭示了与Borda计数一致的明确共识排序。
  • 在Delbeke的家庭构成数据中,该方法检测到两组全局同质群体,分别为68名与14名学生,GHC值分别为75.32%与80.36%,优于传统方法(后者掩盖了较小群体)。
  • 在全局同质群体中,第一TCA因子被证实为Borda计数的仿射函数,验证了该方法的理论基础。
  • 在高维TCA因子中观察到局部异质性,特别是在14名学生的少数群体中,因子得分表现出显著离散与块间交叉。
  • 该方法成功识别出家庭构成偏好中对男孩的偏见,体现在TCA双标图中,(i,j)在第一轴上始终位于(j,i)左侧。
  • GHC系数被证明具有可解释性与有效性,接近100%的值表示块间得分交叉极少,且块内一致性较强。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。