[论文解读] When Computer Vision Gazes at Cognition
本文通过结合心理物理学实验与受人类认知启发的计算机视觉模型,研究了在非受限、三维视觉场景中人类与计算系统在检测第三方视线方向方面的表现。研究发现,人类在区分相距8°–10°的视线目标时准确率约为40%,显著优于当前模型,凸显了人类视线感知的复杂性,也表明需要改进计算框架以更好地模拟这一能力。
Joint attention is a core, early-developing form of social interaction. It is based on our ability to discriminate the third party objects that other people are looking at. While it has been shown that people can accurately determine whether another person is looking directly at them versus away, little is known about human ability to discriminate a third person gaze directed towards objects that are further away, especially in unconstraint cases where the looker can move her head and eyes freely. In this paper we address this question by jointly exploring human psychophysics and a cognitively motivated computer vision model, which can detect the 3D direction of gaze from 2D face images. The synthesis of behavioral study and computer vision yields several interesting discoveries. (1) Human accuracy of discriminating targets 8°-10° of visual angle apart is around 40% in a free looking gaze task; (2) The ability to interpret gaze of different lookers vary dramatically; (3) This variance can be captured by the computational model; (4) Human outperforms the current model significantly. These results collectively show that the acuity of human joint attention is indeed highly impressive, given the computational challenge of the natural looking task. Moreover, the gap between human and model performance, as well as the variability of gaze interpretation across different lookers, require further understanding of the underlying mechanisms utilized by humans for this challenging task.
研究动机与目标
- 理解人类在非受限、自然情境下检测第三方视线方向的视觉感知极限。
- 研究个体在视线解读上的差异如何影响联合注意任务中的表现。
- 开发并评估一种受认知启发的计算机视觉模型,能够从二维人脸图像推断三维视线方向。
- 将人类表现与计算模型进行比较,以识别当前人工智能方法在视线感知方面的差距。
- 揭示在复杂视觉场景中实现高精度视线解读的潜在认知机制。
提出的方法
- 在自由观察视线辨别任务中,对人类受试者开展心理物理学实验,改变潜在视线目标之间的夹角。
- 开发一种受人类认知过程启发的计算机视觉模型,以从二维人脸图像估计三维视线方向。
- 利用三维人脸模型和几何约束,基于单目图像中的眼位与头部朝向推断视线方向。
- 使用人类观察者的数据对模型进行校准,使其预测结果与人类的感知判断保持一致。
- 在具有不同视线解读风格的多个观察者中,评估模型性能与人类数据的一致性。
- 通过区分相距8°–10°视场角的视线目标的准确率来量化性能。
实验结果
研究问题
- RQ1在非受限、自然的视线任务中,人类在区分相距8°–10°的第三方视线目标时的准确率是多少?
- RQ2个体在视线解读上的差异如何影响不同观察者的表现?
- RQ3计算建模系统能否复现人类视线解读中观察到的个体差异?
- RQ4当前计算机视觉模型在视线方向估计方面与人类表现的匹配程度如何,或存在多大差距?
- RQ5在复杂视觉场景中,支撑人类联合注意高敏锐度的潜在认知机制是什么?
主要发现
- 在自由观察视线任务中,人类受试者在区分相距8°–10°视场角的视线目标时,准确率约为40%。
- 不同个体在视线解读能力上表现出显著差异,表明观察者之间的感知策略各不相同。
- 受认知启发的计算机视觉模型成功捕捉到了人类数据中观察到的个体间视线解读差异。
- 尽管如此,模型的性能仍远低于人类水平,表明存在显著的性能差距。
- 结果表明,即使在计算上具有挑战性的自然视觉条件下,人类联合注意的敏锐度依然高度复杂。
- 人类与模型表现之间的差异表明,当前计算机视觉模型在视线感知方面缺乏人类视觉认知中存在的一些关键机制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。