Skip to main content
QUICK REVIEW

[论文解读] Measuring and modeling the perception of natural and unconstrained gaze in humans and machines

Daniel Harari, Tao Gao|arXiv (Cornell University)|Nov 29, 2016
Face Recognition and Perception参考文献 30被引用 7
一句话总结

本文研究了人类与机器在真实世界环境中对自然、非受限注视方向的感知,表明即使在缺乏动态线索的情况下,人类在面对面互动中仍优于机器。通过仅从眼部区域输入学习任务特定表征,一种深度学习模型成功复现了人类的感知模式,包括对沃拉斯顿错觉的敏感性。

ABSTRACT

Humans are remarkably adept at interpreting the gaze direction of other individuals in their surroundings. This skill is at the core of the ability to engage in joint visual attention, which is essential for establishing social interactions. How accurate are humans in determining the gaze direction of others in lifelike scenes, when they can move their heads and eyes freely, and what are the sources of information for the underlying perceptual processes? These questions pose a challenge from both empirical and computational perspectives, due to the complexity of the visual input in real-life situations. Here we measure empirically human accuracy in perceiving the gaze direction of others in lifelike scenes, and study computationally the sources of information and representations underlying this cognitive capacity. We show that humans perform better in face-to-face conditions compared with recorded conditions, and that this advantage is not due to the availability of input dynamics. We further show that humans are still performing well when only the eyes-region is visible, rather than the whole face. We develop a computational model, which replicates the pattern of human performance, including the finding that the eyes-region contains on its own, the required information for estimating both head orientation and direction of gaze. Consistent with neurophysiological findings on task-specific face regions in the brain, the learned computational representations reproduce perceptual effects such as the Wollaston illusion, when trained to estimate direction of gaze, but not when trained to recognize objects or faces.

研究动机与目标

  • 通过实证方法测量人类在逼真、非受限视觉场景中感知注视方向的准确性。
  • 识别支撑人类注视感知的视觉信息来源,例如头部朝向和眼部区域。
  • 开发一种计算模型,以复现人类在注视估计任务中的表现。
  • 探究大脑中的任务特定神经表征是否可在学习到的人工表征中得到映射。
  • 确定动态输入(例如头部运动)是否对人类在注视感知中的优势有贡献。

提出的方法

  • 开展人类心理物理学实验,比较在面对面与录制视频条件下注视感知的准确性。
  • 仅使用人脸的眼部区域数据收集人类表现数据,以隔离其对注视估计的贡献。
  • 训练一种深度卷积神经网络(CNN),从人脸图像中估计注视方向,其架构受分层视觉处理机制的启发。
  • 在其他任务(如人脸分类、物体识别)上训练相同模型,以比较学习到的表征。
  • 通过评估模型在沃拉斯顿错觉等感知错觉上的表现,检验其与人类感知的一致性。
  • 使用激活分析检查模型的内部表征是否与面部处理区域的神经生理学发现相吻合。

实验结果

研究问题

  • RQ1人类在面对面实时互动与录制视频刺激中的注视感知准确性有何差异?
  • RQ2在自然场景中,哪些视觉线索(如头部朝向或眼部区域)对准确的注视估计贡献最大?
  • RQ3在注视估计任务上训练的深度学习模型能否复现人类的感知模式,包括对视觉错觉的敏感性?
  • RQ4为注视估计学习到的表征与为人脸或物体识别学习到的表征有何不同?它们是否反映了任务特定的神经组织结构?
  • RQ5人类在注视感知中的优势是源于动态输入,还是其他因素(如情境或社会线索)?

主要发现

  • 即使在去除动态线索后,人类在面对面注视感知中的表现仍显著优于录制视频条件。
  • 仅眼部区域就包含足够信息以实现准确的注视估计,且当仅可见该区域时,人类表现依然很高。
  • 在注视方向估计任务上训练的深度学习模型复现了人类的表现模式,包括对沃拉斯顿错觉的敏感性。
  • 该模型的内部表征与神经生理学发现相符:它发展出针对注视估计的任务特定特征,但在训练为人脸或物体识别时则未出现。
  • 人类在注视感知中的优势并非源于输入动态性,表明其依赖于更高级别的情境或社会处理。
  • 该模型在复现感知错觉方面的成功表明,其内部表征捕捉了人类视觉处理中与注视感知相关的关键方面。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。