Skip to main content
QUICK REVIEW

[论文解读] A Novel Space-Time Representation on the Positive Semidefinite Con for Facial Expression Recognition

Anis Kacem, Mohamed Daoudi|arXiv (Cornell University)|Jul 20, 2017
Face and Expression Recognition参考文献 34被引用 11
一句话总结

本文通过将面部关键点的时间序列建模为固定秩正半定矩阵的黎曼流形上的轨迹,提出了一种新颖的时空表示方法用于面部表情识别。通过结合仿射不变形状与空间协方差信息,该方法利用几何工具(如时间对齐和自适应重采样)实现了速率不变的轨迹分析,在 CK+、MMI、Oulu-CASIA 和 AFEW 数据集上实现了最先进性能,CK+ 上准确率最高达 96.87%,AFEW 上达 39.94%。

ABSTRACT

In this paper, we study the problem of facial expression recognition using a novel space-time geometric representation. We describe the temporal evolution of facial landmarks as parametrized trajectories on the Riemannian manifold of positive semidefinite matrices of fixed-rank. Our representation has the advantage to bring naturally a second desirable quantity when comparing shapes -- the spatial covariance -- in addition to the conventional affine-shape representation. We derive then geometric and computational tools for rate-invariant analysis and adaptive re-sampling of trajectories, grounding on the Riemannian geometry of the manifold. Specifically, our approach involves three steps: 1) facial landmarks are first mapped into the Riemannian manifold of positive semidefinite matrices of rank 2, to build time-parameterized trajectories; 2) a temporal alignment is performed on the trajectories, providing a geometry-aware (dis-)similarity measure between them; 3) finally, pairwise proximity function SVM (ppfSVM) is used to classify them, incorporating the latter (dis-)similarity measure into the kernel function. We show the effectiveness of the proposed approach on four publicly available benchmarks (CK+, MMI, Oulu-CASIA, and AFEW). The results of the proposed approach are comparable to or better than the state-of-the-art methods when involving only facial landmarks.

研究动机与目标

  • 解决使用对刚性变换不变的几何表示来建模动态面部表情的挑战。
  • 将空间协方差整合到形状表示中,以超越标准仿射形状分析的判别能力。
  • 利用黎曼几何实现面部表情轨迹的速率不变比较与分类。
  • 在正半定锥面上开发用于关键点轨迹时间对齐与自适应重采样的计算工具。
  • 仅使用面部关键点,在多个基准数据集上实现最先进性能。

提出的方法

  • 通过格拉姆矩阵将 2D 面部关键点配置映射到秩为 2 的正半定矩阵的黎曼流形上。
  • 将面部表情的时间演化表示为在流形 $\mathcal{S}^{+}(2,n)$ 上的时间参数化轨迹。
  • 应用动态时间规整(DTW)对轨迹进行速率不变的时间对齐,以确保几何一致性。
  • 使用定义在流形上的接近度度量 $d_{\mathcal{S}^{+}}$ 计算轨迹之间的(非)相似性。
  • 实施自适应重采样以标准化轨迹长度并提高分类鲁棒性。
  • 采用基于 $d_{\mathcal{S}^{+}}$ 距离的核函数的成对邻近函数支持向量机(ppfSVM)进行有效分类。

实验结果

研究问题

  • RQ1如何在黎曼框架下对几何方式表示面部关键点轨迹,以同时保留形状与空间协方差信息?
  • RQ2在非欧几里得流形上,以何种方式最有效地实现动态面部表情序列的速率不变比较与对齐?
  • RQ3将空间协方差整合到形状表示中,相较于仅使用仿射形状的方法,如何提升面部表情识别性能?
  • RQ4时间对齐与自适应重采样等几何工具在基准数据集上的分类准确率提升程度如何?
  • RQ5ppfSVM 分类器能否有效学习正半定锥面上的非欧几里得轨迹表示?

主要发现

  • 所提出的 $d_{\mathcal{S}^{+}}$ 距离在 $\mathcal{S}^{+}(2,n)$ 上优于平坦的弗罗贝尼乌斯距离与 $d_{\mathcal{P}_n}$ 距离,在 CK+ 上达到 96.87% 的准确率。
  • 通过 DTW 实现的时间对齐使 CK+ 上准确率提升约 6%,MMI 上提升约 12%,证明其在速率不变分析中的关键作用。
  • 自适应重采样使 MMI 上性能提升约 5%,AFEW 上提升约 3%,增强了轨迹的一致性以利于分类。
  • ppfSVM 分类器在 AFEW 上达到 39.94% 的准确率,在仅使用面部关键点的方法中排名第二,优于使用多数投票的 K-NN(29.77%)。
  • 该方法在四个基准数据集(CK+、MMI、Oulu-CASIA 和 AFEW)上实现了最先进或具有竞争力的结果,且未依赖外观特征。
  • 与基于格拉斯曼流形的方法相比,该框架的优越性凸显了通过正半定表示引入空间协方差所带来的额外判别能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。