Skip to main content
QUICK REVIEW

[论文解读] Matching Representations of Explainable Artificial Intelligence and Eye Gaze for Human-Machine Interaction

Tiffany Hwu, Mia Levy|arXiv (Cornell University)|Jan 30, 2021
Visual Attention and Saliency Detection参考文献 17被引用 9
一句话总结

本文提出在驾驶场景中通过逐层相关性传播(LRP)对可解释人工智能(XAI)进行对齐,使其与人类眼球注视相匹配,从而实现非语言的人机通信。结果表明,随着任务特异性的提高,驾驶预测神经网络的LRP热力图与人类眼球注视的相似性逐渐增强,揭示了无需显式目标标注的共享注意力表征。

ABSTRACT

Rapid non-verbal communication of task-based stimuli is a challenge in human-machine teaming, particularly in closed-loop interactions such as driving. To achieve this, we must understand the representations of information for both the human and machine, and determine a basis for bridging these representations. Techniques of explainable artificial intelligence (XAI) such as layer-wise relevance propagation (LRP) provide visual heatmap explanations for high-dimensional machine learning techniques such as deep neural networks. On the side of human cognition, visual attention is driven by the bottom-up and top-down processing of sensory input related to the current task. Since both XAI and human cognition should focus on task-related stimuli, there may be overlaps between their representations of visual attention, potentially providing a means of nonverbal communication between the human and machine. In this work, we examine the correlations between LRP heatmap explanations of a neural network trained to predict driving behavior and eye gaze heatmaps of human drivers. The analysis is used to determine the feasibility of using such a technique for enhancing driving performance. We find that LRP heatmaps show increasing levels of similarity with eye gaze according to the task specificity of the neural network. We then propose how these findings may assist humans by visually directing attention towards relevant areas. To our knowledge, our work provides the first known analysis of LRP and eye gaze for driving tasks.

研究动机与目标

  • 探究可解释人工智能(XAI)表征(特别是LRP热力图)是否与驾驶过程中的人类视觉注意力对齐。
  • 确定LRP是否能在无需显式目标检测或分割的情况下识别与任务相关的视觉刺激。
  • 探索基于LRP的视觉提示在人机协同系统中引导人类注意力的可行性。
  • 评估LRP与眼球注视作为自适应驾驶员辅助系统中互补信号的潜力。

提出的方法

  • 在DR(eye)VE数据集上训练了一个基于VGG16的深度神经网络,以从视频帧中预测驾驶行为。
  • 应用逐层相关性传播(LRP)为每帧生成显著性热力图,突出对预测贡献最大的图像区域。
  • 使用眼动追踪眼镜收集人类驾驶员在自然驾驶任务中进行眼动追踪的数据。
  • 从驾驶员眼动追踪数据中生成注视热力图,并与LRP热力图在空间和统计上进行比较。
  • 采用减法分析方法,分离出驾驶专用LRP模型强调但ImageNet预训练模型未强调的特征。
  • 使用相关性度量评估LRP与注视热力图之间的相似性,重点关注任务特定与一般视觉显著性之间的差异。

实验结果

研究问题

  • RQ1驾驶预测神经网络的LRP热力图在真实驾驶场景中与人类眼球注视的相关性有多大?
  • RQ2神经网络训练中的任务特异性如何影响LRP显著性与人类视觉注意力之间的对齐程度?
  • RQ3LRP是否能在无显式目标标注的情况下识别与任务相关的视觉特征(如车道线或交通标志)?
  • RQ4通过减法分析,通用ImageNet模型与专用驾驶模型在显著性模式上存在哪些差异?

主要发现

  • 驾驶专用神经网络生成的LRP热力图与人类眼球注视的空间相关性显著高于通用ImageNet模型生成的热力图。
  • 随着任务特异性的提高,LRP与眼球注视之间的相似性增强,表明LRP即使在缺乏显式目标级监督的情况下也能捕捉与任务相关的特征。
  • 减法分析显示,驾驶专用LRP模型强调了车道线、边界和道路特征,而ImageNet模型未突出这些内容,表明其可能实现了对任务相关刺激的潜在发现。
  • LRP热力图成功突出了停车标志、交通信号灯和行人等显著道路元素,即使这些在训练数据中未被显式标注。
  • 该方法在过滤无关视觉刺激方面表现出潜力,支持在驾驶员辅助系统中设计非侵入式、注意力引导的视觉提示。
  • 结果表明,LRP可作为闭环人机交互中人类注意力的代理,尤其在驾驶等安全关键领域具有应用前景。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。