Skip to main content
QUICK REVIEW

[论文解读] Limitations and Biases in Facial Landmark Detection -- An Empirical Study on Older Adults with Dementia

Azin Asgarian, Shun Zhao|arXiv (Cornell University)|May 17, 2019
Face recognition and analysis参考文献 31被引用 8
一句话总结

本研究调查了针对痴呆症老年群体的面部关键点检测中的算法偏见,评估了七种最先进方法在正面和侧面对脸图像上的表现。尽管在六个数据集上进行了再训练和微调,性能差距依然存在,尤其是在口部、眼部和鼻部区域,表明存在固有偏见,从而损害了痴呆症评估中临床应用的可靠性。

ABSTRACT

Accurate facial expression analysis is an essential step in various clinical applications that involve physical and mental health assessments of older adults (e.g. diagnosis of pain or depression). Although remarkable progress has been achieved toward developing robust facial landmark detection methods, state-of-the-art methods still face many challenges when encountering uncontrolled environments, different ranges of facial expressions, and different demographics of the population. A recent study has revealed that the health status of individuals can also affect the performance of facial landmark detection methods on front views of faces. In this work, we investigate this matter in a much greater context using seven facial landmark detection methods. We perform our evaluation not only on frontal faces but also on profile faces and in various regions of the face. Our results shed light on limitations of the existing methods and challenges of applying these methods in clinical settings by indicating: 1) a significant difference between the performance of state-of-the-art when tested on the profile or frontal faces of individuals with vs. without dementia; 2) insights on the existing bias for all regions of the face; and 3) the presence of this bias despite re-training/fine-tuning with various configurations of six datasets.

研究动机与目标

  • 调查面部关键点检测方法在应用于痴呆症老年群体与认知健康对照组时是否存在偏见。
  • 评估最先进关键点检测模型在不同面部区域的正面和侧面对脸图像上的表现。
  • 评估使用多样化数据集进行再训练或微调是否能缓解痴呆症群体与健康对照组之间的性能差异。
  • 识别在痴呆症个体中检测精度显著下降的具体面部区域。
  • 为针对老龄化和神经退行性病变人群的临床面部分析系统中的算法偏见提供实证证据。

提出的方法

  • 评估了七种面部关键点检测方法:AAM、CFSS、CLNF、FAN-2D、FAN-3D、PRNet 和 Mnemonic Descent。
  • 使用六个基准数据集:Helen、AFW、LFPW、MENPO Profile、UNBC-McMaster 痛苦档案和痴呆症痛苦数据集。
  • 在痴呆症痛苦数据集的正面(Tf)和侧面对脸(Tp)子集上进行实验,比较健康组与痴呆组的表现。
  • 采用多种训练配置进行再训练和微调:S1、S2、S1∪S2、Tf∪S1、Tf∪S2 和 Tf∪S1∪S2。
  • 使用收敛曲线和在5%容差范围内的RMS拟合误差测量性能,并进行显著性检验(p值)。
  • 分析下颌、眉毛、鼻部、眼部和口部等区域的性能,以识别偏见热点。

实验结果

研究问题

  • RQ1痴呆症老年群体与认知健康老年群体在面部关键点检测表现上是否存在显著差异?
  • RQ2在某些面部区域(如口部、眼部、鼻部)中,关键点检测的性能差距是否比其他区域更明显?
  • RQ3使用多样化数据集进行再训练或微调在多大程度上可以减少痴呆症群体与健康对照组之间的性能差距?
  • RQ4在痴呆症个体中,关键点检测模型在正面和侧面对脸视图下的表现如何变化?
  • RQ5尽管架构和训练策略各异,这种观察到的偏见是否在多种最先进方法中保持一致?

主要发现

  • 在正面脸图像上,痴呆症老年群体与健康对照组之间存在显著性能差距,尤其在口部(52.66% vs. 37.54% 收敛率)、眼部(47.93% vs. 44.00%)和鼻部(51.18% vs. 46.15%)区域。
  • 即使在使用组合数据集进行再训练后,性能差距依然具有统计显著性(p < 0.001),表明存在持续的算法偏见。
  • 整体而言,侧面对脸检测性能较低,但痴呆症组与健康组之间的差距小于正面脸,尽管在鼻部和口部区域的差距仍然显著。
  • 再训练提高了两组的收敛率,但所有方法和配置下,痴呆症个体与健康个体之间的相对性能差距依然存在。
  • AAM 方法在正面脸上的整体表现最高(健康子集收敛率为44.67%),但在痴呆症脸上的表现仍较低(收敛率为37.23%)。
  • 在侧面对脸中,PRNet 在健康子集上达到最高收敛率(41.12%),但在痴呆症子集上仅为28.92%,整个面部区域的p值 < 0.001。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。