[论文解读] A Study on Multimodal and Interactive Explanations for Visual Question Answering
本文提出了一种多模态、交互式的视觉问答(VQA)解释系统,结合视觉注意力图、文本解释以及一种允许用户编辑注意力图的交互式“主动注意力”界面。通过90名参与者的用户研究,结果表明,解释显著提升了用户预测VQA系统准确性的能力,尤其是在系统出错时;同时,解释评分与预测表现高度相关,验证了其在人机协作中的有效性。
Explainability and interpretability of AI models is an essential factor affecting the safety of AI. While various explainable AI (XAI) approaches aim at mitigating the lack of transparency in deep networks, the evidence of the effectiveness of these approaches in improving usability, trust, and understanding of AI systems are still missing. We evaluate multimodal explanations in the setting of a Visual Question Answering (VQA) task, by asking users to predict the response accuracy of a VQA agent with and without explanations. We use between-subjects and within-subjects experiments to probe explanation effectiveness in terms of improving user prediction accuracy, confidence, and reliance, among other factors. The results indicate that the explanations help improve human prediction accuracy, especially in trials when the VQA system's answer is inaccurate. Furthermore, we introduce active attention, a novel method for evaluating causal attentional effects through intervention by editing attention maps. User explanation ratings are strongly correlated with human prediction accuracy and suggest the efficacy of these explanations in human-machine AI collaboration tasks.
研究动机与目标
- 评估多模态与交互式解释是否能提升人类对VQA系统行为的理解与预测能力。
- 研究解释如何影响人类在人机协作中的信心、依赖程度及心智模型的构建。
- 提出并评估一种新型交互方法——‘主动注意力’,允许用户修改注意力图并实时观察系统反馈。
- 衡量解释有用性评分与用户预测准确率之间的相关性。
- 评估不同解释类型对用户表现的影响,特别是在VQA系统给出错误答案时。
提出的方法
- 通过90名参与者的组间与组内用户研究,比较在有无解释情况下的预测准确率。
- 提供多模态解释,包括空间注意力图、文本解释,以及一种新颖的交互式‘主动注意力’界面以编辑注意力图。
- 收集用户对每项预测任务中解释有用性及信心水平的评分。
- 采用正确性预测任务,用户预测VQA系统答案是否正确,性能在有解释与无解释的区块中进行衡量。
- 通过允许用户修改注意力图并即时观察反馈,实施主动注意力,从而对注意力影响进行因果评估。
- 分析解释评分、预测准确率与用户信心之间的相关性,以评估解释的有效性。
实验结果
研究问题
- RQ1提供多模态解释(视觉、文本及交互式)是否能提升用户预测VQA系统答案正确性的能力?
- RQ2解释评分与用户在VQA任务中的预测准确率及信心水平之间有何相关性?
- RQ3交互式解释(主动注意力)对用户在预测VQA结果时的信任度与表现有何影响?
- RQ4当VQA系统正确或错误时,用户预测行为有何差异,尤其在使用解释时?
- RQ5当解释模式存在冲突或信息过载时,用户表现会受到多大程度的负面影响?
主要发现
- 解释显著提升了用户预测准确率,尤其在VQA系统给出错误答案的测试中表现更明显。
- 将解释评为更有用的用户,更有可能准确预测系统的正确性,表明解释质量与用户表现之间存在强烈相关性。
- 用户对预测的信心与VQA系统自身的置信度分数(最高答案概率)高度相关,表明用户已建立起对系统的可靠心智模型。
- 与其它解释模式相比,主动注意力界面显著提升了用户信心,表明其在建立用户信任与理解方面的有效性。
- 在暴露于多种冲突解释模式的组别(如AL组)中,表现未见提升,部分用户报告因信息过载或信号冲突而感到困惑。
- 在系统出错时,被评价为有用的解释反而具有误导性,凸显了在系统失败场景下设计更稳健解释机制的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。