Skip to main content
QUICK REVIEW

[论文解读] Tensor Analysis with n-Mode Generalized Difference Subspace

Bernardo B. Gatto, Eulanda M. dos Santos|arXiv (Cornell University)|Sep 4, 2019
Human Pose and Action Recognition参考文献 64被引用 4
一句话总结

本文提出了一种新型基于张量的分类方法——n模广义差异子空间(n-mode GDS),该方法利用多线性代数和在乘积Grassmann流形上的判别子空间学习。通过整合n模奇异值分解、用于衡量模重要性的n模Fisher得分,以及加权测地距离度量,该方法在无需预训练模型或迁移学习的情况下,在动作识别和手势识别任务中达到了最先进性能。

ABSTRACT

The increasing use of multiple sensors, which produce a large amount of multi-dimensional data, requires efficient representation and classification methods. In this paper, we present a new method for multi-dimensional data classification that relies on two premises: 1) multi-dimensional data are usually represented by tensors, since this brings benefits from multilinear algebra and established tensor factorization methods; and 2) multilinear data can be described by a subspace of a vector space. The subspace representation has been employed for pattern-set recognition, and its tensor representation counterpart is also available in the literature. However, traditional methods do not use discriminative information of the tensors, degrading the classification accuracy. In this case, generalized difference subspace (GDS) provides an enhanced subspace representation by reducing data redundancy and revealing discriminative structures. Since GDS does not handle tensor data, we propose a new projection called n-mode GDS, which efficiently handles tensor data. We also introduce the n-mode Fisher score as a class separability index and an improved metric based on the geodesic distance for tensor data similarity. The experimental results on gesture and action recognition show that the proposed method outperforms methods commonly used in the literature without relying on pre-trained models or transfer learning.

研究动机与目标

  • 解决传统张量分类方法在多维数据中未能有效利用判别结构的局限性。
  • 开发一种保持高阶数据时空关系并减少冗余的张量分类框架。
  • 提出一种专为在乘积Grassmann流形上的张量数据设计的判别性、几何感知子空间学习方法。
  • 通过使用手工设计特征和张量内在结构,消除对预训练深度神经网络的依赖,以提升泛化能力。
  • 提供一种在张量模数上具有线性复杂度的可扩展解决方案,适用于小样本、高维数据集。

提出的方法

  • 提出n-mode GDS,一种投影方法,将广义差异子空间(GDS)应用于张量的每个模,实现在多个模上的判别性表征。
  • 引入n模Fisher得分作为模特定的可分性指标,以自动加权各张量模的重要性。
  • 在乘积Grassmann流形(PGM)上采用加权测地距离度量,以在保持几何结构的同时测量张量子空间之间的相似性。
  • 通过张量展开(矩阵化)分别分析每个模,实现模级处理和子空间学习。
  • 将手工设计特征(HOG、HOF、MBH)与基于张量的表征相结合,以增强判别能力,且无需预训练。
  • 应用n模奇异值分解(n-SVD)从每个张量模中提取低秩表征,以减少冗余并保留内在结构。

实验结果

研究问题

  • RQ1广义差异子空间方法能否被有效扩展至多模张量数据,以提升分类准确率?
  • RQ2如何最优地结合多模张量中的判别信息,以提升类别可分性?
  • RQ3所提出的n模Fisher得分是否能提供一种可靠且自动的手段,用于估计不同张量模的相对重要性?
  • RQ4在乘积Grassmann流形上使用加权测地距离能否增强张量分类中的相似性度量?
  • RQ5在不依赖深度学习或迁移学习的情况下,手工设计特征能在多大程度上提升张量分类性能?

主要发现

  • n-mode GDS方法在动作识别和手势识别任务中优于C3D和双流CNN等最先进方法,且在无需预训练模型的情况下实现了具有竞争力的准确率。
  • 在KTH和UCF-101数据集上,使用HOG特征使n-mode GDS准确率分别提升了约2%,证实了外观特征的优势。
  • 引入HOF和MBH特征后,准确率分别提升了约7%,表明在所提框架中时间特征具有显著优势。
  • 基于n模Fisher得分的加权策略在两个数据集上分别使性能提升了1.5%和2%,通过自动为更具判别性的模分配更高权重。
  • 当结合所有描述符(HOG、HOF、MBH)时,方法达到最优性能,其中模式1、2、3的模权重分别为0.32、0.29和0.38(使用HOG特征时)。
  • 结果证实,在流形上解释时空模并平衡其贡献,对于最大化分类准确率至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。