Skip to main content
QUICK REVIEW

[论文解读] Generalized Categorization Axioms

Yu Jian|arXiv (Cornell University)|Mar 31, 2015
Face and Expression Recognition参考文献 19被引用 4
一句话总结

本文提出了一种广义分类公理框架,通过内类别表示与外类别表示重新诠释分类问题,将聚类、分类、回归、降维和密度估计统一于基于相似性的公理之下。该框架通过引入类别表示公理,克服了以往分类公理在类别数为1时的局限性,确保了在各类机器学习任务中,尤其是当类别数为1时的理论一致性。

ABSTRACT

Categorization axioms have been proposed to axiomatizing clustering results, which offers a hint of bridging the difference between human recognition system and machine learning through an intuitive observation: an object should be assigned to its most similar category. However, categorization axioms cannot be generalized into a general machine learning system as categorization axioms become trivial when the number of categories becomes one. In order to generalize categorization axioms into general cases, categorization input and categorization output are reinterpreted by inner and outer category representation. According to the categorization reinterpretation, two category representation axioms are presented. Category representation axioms and categorization axioms can be combined into a generalized categorization axiomatic framework, which accurately delimit the theoretical categorization constraints and overcome the shortcoming of categorization axioms. The proposed axiomatic framework not only discuses categorization test issue but also reinterprets many results in machine learning in a unified way, such as dimensionality reduction,density estimation, regression, clustering and classification.

研究动机与目标

  • 为解决现有分类公理在类别数为1时变得平凡的问题,通过内类别表示与外类别表示重新诠释分类的输入与输出。
  • 建立一个统一的理论框架,将分类公理推广至所有机器学习任务,包括聚类、分类、回归和降维。
  • 形式化相似性在类人化分类中的作用,并通过公理约束将认知科学原理与机器学习相连接。
  • 提供基于一致性、鲁棒性和表示保真度的理论标准,用于评估与设计分类算法。

提出的方法

  • 提出双表示模型:外类别表示(显式隶属关系)与内类别表示(隐式相似性),二者均定义于输入与输出空间。
  • 提出两类类别表示公理:存在性与唯一性,确保类别表示在变换下具有良好的定义性与稳定性。
  • 在新表示框架下重新诠释核心分类公理——等价性、一致性、鲁棒性与误差,以确保所有学习任务中的理论有效性。
  • 基于相似性分配公理(SS)定义关键概念,如决策区域、训练决策区域与间隔,实现分类边界几何化解释。
  • 应用分类一致性原理(CCP)与分类鲁棒性假设(CRA),推导学习算法的理论稳定性条件。
  • 将既有的算法(如PCA、SVM、Lasso、LDA)重新诠释为该框架的特例,表明在适当的内类别表示下,它们均满足公理。

实验结果

研究问题

  • RQ1如何将分类公理推广至类别数为1的情况,例如在降维或回归任务中?
  • RQ2在输入数据存在噪声或部分缺失时,何种条件可确保分类算法的理论稳定性和一致性?
  • RQ3内类别表示与外类别表示如何共同约束机器学习模型的设计与评估?
  • RQ4相似性在单一公理框架下统一各类学习任务中扮演何种角色?
  • RQ5如何形式化并应用认知科学中的类别表示原则,以改进机器学习算法?

主要发现

  • 所提出的广义分类公理框架通过引入内类别表示与外类别表示,克服了以往分类公理在c=1时的平凡性问题。
  • 类别表示公理——存在性与唯一性——确保了类别表示的良定义性与稳定性,为学习算法奠定了坚实的理论基础。
  • 该框架将降维、密度估计、回归、聚类与分类统一于单一公理结构之下,表明它们均为广义分类的特例。
  • 当c=1时,该框架解释了为何PCA、NMF、Lasso与Isomap等方法满足分类一致性原理(CCP),而CCP等价于经验风险最小化或结构风险最小化。
  • 对于分类任务(c>1),基于相似性的分配公理(SS)与一致性公理(CS、UCR)提供了理论约束,可解释SVM、LDA与逻辑回归等模型的行为。
  • 分类鲁棒性假设(CRA)为算法稳定性提供了理论条件,尤其在输入数据存在噪声或部分观测时表现显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。