[论文解读] Towards Reconciling Usability and Usefulness of Explainable AI Methodologies
本研究评估了四种可解释人工智能(XAI)模态——决策树、自然语言文本、可执行程序和注意力图——在高速公路模拟环境中解释自动驾驶汽车决策的效果。研究发现一个关键矛盾:尽管用户主观上更偏好文本解释以提升可用性,但决策树在客观上更有利于预测汽车行为,凸显了在个性化XAI设计中平衡感知可用性与实际有用性的重要性。
Interactive Artificial Intelligence (AI) agents are becoming increasingly prevalent in society. However, application of such systems without understanding them can be problematic. Black-box AI systems can lead to liability and accountability issues when they produce an incorrect decision. Explainable AI (XAI) seeks to bridge the knowledge gap, between developers and end-users, by offering insights into how an AI algorithm functions. Many modern algorithms focus on making the AI model "transparent", i.e. unveil the inherent functionality of the agent in a simpler format. However, these approaches do not cater to end-users of these systems, as users may not possess the requisite knowledge to understand these explanations in a reasonable amount of time. Therefore, to be able to develop suitable XAI methods, we need to understand the factors which influence subjective perception and objective usability. In this paper, we present a novel user-study which studies four differing XAI modalities commonly employed in prior work for explaining AI behavior, i.e. Decision Trees, Text, Programs. We study these XAI modalities in the context of explaining the actions of a self-driving car on a highway, as driving is an easily understandable real-world task and self-driving cars is a keen area of interest within the AI community. Our findings highlight internal consistency issues wherein participants perceived language explanations to be significantly more usable, however participants were better able to objectively understand the decision making process of the car through a decision tree explanation. Our work also provides further evidence of importance of integrating user-specific and situational criteria into the design of XAI systems. Our findings show that factors such as computer science experience, and watching the car succeed or fail can impact the perception and usefulness of the explanation.
研究动机与目标
- 探究不同XAI模态如何影响用户对自主系统中AI决策的感知与客观理解。
- 考察用户特定因素(如计算机科学经验)以及情境因素(例如观察汽车成功或失败)对XAI偏好与有效性的影响力。
- 识别XAI解释中感知可用性与实际有用性之间的脱节现象。
- 为设计能够根据个体用户认知与情境特征自适应调整的个性化XAI系统提供依据。
提出的方法
- 在包含231名参与者的大型用户研究中,使用模拟自动驾驶汽车环境,评估四种XAI模态:决策树、自然语言文本、可执行程序和注意力图。
- 采用前向仿真协议,参与者在查看解释后预测汽车行为,从而实现对理解程度的客观测量。
- 通过问卷调查收集主观可用性评分,以评估各模态的感知帮助程度与信任度。
- 分析人口统计因素(如计算机科学背景)与情境因素(如观察汽车成功或发生碰撞)对解释感知与表现的影响。
- 采用混合方法分析,结合定量预测准确率与定性反馈,以评估解释质量。
- 提出一个未来个性化XAI系统的框架,该框架可根据用户倾向与情境动态调整解释方式。
实验结果
研究问题
- RQ1不同XAI模态(决策树、文本、程序、注意力图)在感知可用性与对AI决策的客观理解方面如何比较?
- RQ2计算机科学经验等个体因素如何影响用户对XAI解释的感知与表现?
- RQ3情境因素(特别是观察自动驾驶汽车成功或失败)如何影响用户对XAI解释的信任与评价?
- RQ4感知可用性与XAI解释实际有用性之间存在多大程度的脱节?
- RQ5能否设计一种个性化XAI系统,根据用户特征与情境动态选择最合适的解释模态?
主要发现
- 基于文本的解释在主观可用性调查中被参与者普遍评价为显著更易用,优于其他模态。
- 尽管感知可用性更高,但基于文本的解释在预测汽车行为方面的客观准确率低于决策树。
- 决策树在客观理解方面表现最佳,参与者在预测准确率上显著优于基于文本或程序的解释。
- 计算机科学经验更高的参与者对基于文本的解释偏好降低,表明先前知识可能改变模态偏好。
- 观察到汽车发生碰撞(失败)显著降低了用户对XAI代理的态度与信任度,尤其在基于文本的解释中更为明显。
- 本研究揭示了一个根本性矛盾:感知可用性(主观)与实际有用性(客观)不一致,即文本解释虽更受青睐,但在准确预测方面效果更差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。