[论文解读] Disentangled Variational Representation for Heterogeneous Face Recognition
本文提出了一种解耦变分表征(DVR)框架,用于异质性人脸验证,通过变分推断在近红外(NIR)和可见光(VIS)人脸表征中解耦身份信息与个体内变化。通过最小化同一主体的身份信息,并在模态差异上应用松弛相关对齐,该方法减少了域差距并缓解了过拟合,在三个基准数据集上实现了最先进性能,尤其在Oulu-CASIA数据集上VR@FAR=1%的性能提升高达19.1%。
Visible (VIS) to near infrared (NIR) face matching is a challenging problem due to the significant domain discrepancy between the domains and a lack of sufficient data for training cross-modal matching algorithms. Existing approaches attempt to tackle this problem by either synthesizing visible faces from NIR faces, extracting domain-invariant features from these modalities, or projecting heterogeneous data onto a common latent space for cross-modal matching. In this paper, we take a different approach in which we make use of the Disentangled Variational Representation (DVR) for cross-modal matching. First, we model a face representation with an intrinsic identity information and its within-person variations. By exploring the disentangled latent variable space, a variational lower bound is employed to optimize the approximate posterior for NIR and VIS representations. Second, aiming at obtaining more compact and discriminative disentangled latent space, we impose a minimization of the identity information for the same subject and a relaxed correlation alignment constraint between the NIR and VIS modality variations. An alternative optimization scheme is proposed for the disentangled variational representation part and the heterogeneous face recognition network part. The mutual promotion between these two parts effectively reduces the NIR and VIS domain discrepancy and alleviates over-fitting. Extensive experiments on three challenging NIR-VIS heterogeneous face recognition databases demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods.
研究动机与目标
- 解决异质性人脸识别中近红外(NIR)与可见光(VIS)人脸图像之间显著的域差距问题。
- 减少因近红外(NIR)人脸识别训练数据有限而导致的过拟合问题。
- 学习紧凑、判别性强且解耦的潜在表征,将身份信息与个体内变化分离。
- 通过联合优化解耦表征与识别网络,提升跨模态匹配性能。
- 开发解耦变分表征与HFR网络之间的相互优化机制,以增强泛化能力。
提出的方法
- 采用变分自编码器框架建模人脸表征,将身份与个体内变化解耦为独立的潜在变量。
- 使用变分下界优化NIR与VIS模态的近似后验分布。
- 最小化同一主体的身份信息,以增强解耦潜在空间中的判别能力。
- 在模态差异之间应用松弛相关对齐约束,以对齐光照与光谱差异。
- 引入一种交替优化方案,交替更新DVR模块与HFR网络,实现相互改进。
- 从学习到的后验分布生成合成的NIR与VIS样本,用于数据增强,减少过拟合。
实验结果
研究问题
- RQ1解耦变分表征能否在异质性人脸识别中有效分离身份与个体内变化?
- RQ2最小化同一主体的身份信息在多大程度上能提升跨模态匹配性能?
- RQ3模态差异之间的松弛相关对齐在多大程度上能减少域差距?
- RQ4解耦表征与HFR网络之间的相互优化能否缓解小规模NIR数据集上的过拟合问题?
- RQ5所提出的DVR框架是否在具有挑战性的NIR-VIS人脸识别基准上优于最先进方法?
主要发现
- 在CASIA NIR-VIS 2.0数据集上,DVR结合LightCNN-9实现99.1%的Rank-1准确率和98.6%的VR@FAR=0.1%,优于所有对比的最先进方法。
- 使用LightCNN-29时,DVR相较LightCNN-9将Rank-1准确率提升0.8%,VR@FAR=0.1%提升1.0%,表明其在更深网络中的可扩展性。
- 在Oulu-CASIA NIR-VIS数据集上,DVR实现89.7%的VR@FAR=1%,相较之前SOTA方法W-CNN(81.5%)提升8.2%。
- 在BUAA-VisNir数据集上,DVR实现97.0%的VR@FAR=1%,优于W-CNN 1.0%,展现出在小样本数据上的强泛化能力。
- 消融实验表明,每个组件——解耦变分部分、均值差异最小化和相关对齐——均对性能提升有显著贡献。
- ROC曲线显示,DVR始终优于所有基线方法,尤其在低假阳性率下表现更优,表明其具备鲁棒性和高判别能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。