[论文解读] X-ToM: Explaining with Theory-of-Mind for Gaining Justified Human Trust
X-ToM 提出了一种基于心智理论的解释框架,将AI系统的行为建模为具有心理状态,使用户能够理解预测的*原因*。与问答(QA)和显著性图基线相比,该方法在解释满意度方面有显著提升(p<0.01),尤其在实用性、充分性和细节方面表现更优,且未增加响应时间,从而促进人类对AI的合理信任。
We present a new explainable AI (XAI) framework aimed at increasing justified human trust and reliance in the AI machine through explanations. We pose explanation as an iterative communication process, i.e. dialog, between the machine and human user. More concretely, the machine generates sequence of explanations in a dialog which takes into account three important aspects at each dialog turn: (a) human's intention (or curiosity); (b) human's understanding of the machine; and (c) machine's understanding of the human user. To do this, we use Theory of Mind (ToM) which helps us in explicitly modeling human's intention, machine's mind as inferred by the human as well as human's mind as inferred by the machine. In other words, these explicit mental representations in ToM are incorporated to learn an optimal explanation policy that takes into account human's perception and beliefs. Furthermore, we also show that ToM facilitates in quantitatively measuring justified human trust in the machine by comparing all the three mental representations. We applied our framework to three visual recognition tasks, namely, image classification, action recognition, and human body pose estimation. We argue that our ToM based explanations are practical and more natural for both expert and non-expert users to understand the internal workings of complex machine learning models. To the best of our knowledge, this is the first work to derive explanations using ToM. Extensive human study experiments verify our hypotheses, showing that the proposed explanations significantly outperform the state-of-the-art XAI methods in terms of all the standard quantitative and qualitative XAI evaluation metrics including human trust, reliance, and explanation satisfaction.
研究动机与目标
- 为弥合人类对AI信任的缺口,开发一种将AI推理过程建模为具有心理状态的解释方法。
- 超越如显著性图或问答等表面化解释,提升用户对AI预测的理解。
- 衡量心智理论解释是否能带来更高的合理信任度与依赖度。
- 在受控的人工参与研究中,评估不同解释类型在用户满意度、响应时间与依赖度方面的表现。
提出的方法
- X-ToM 框架通过将AI建模为具有信念与意图的‘表演者’,模拟其如何推断物体部件及其关系,从而生成解释。
- 采用类似场景图的结构(AOG)来编码物体部件及其关系,支持对最具影响力的图像区域进行推理。
- 针对检测成功原因及关键图像区域等评估者问题生成解释,例如:'哪些部分最有助于识别‘跑步的人’?'
- 通过网络界面开展评估,共招募120名来自心理学人群池的参与者,将X-ToM与QA和显著性图基线进行比较。
- 用户通过李克特量表对满意度、依赖度与响应时间进行评分,以评估可用性与信任度。
- 评估结合定性与定量指标,包括解释满意度、响应时间及主观依赖度。
实验结果
研究问题
- RQ1与问答(QA)和显著性图相比,X-ToM是否提升了用户对解释的满意度?
- RQ2用户是否比基线方法更快地理解X-ToM的解释?
- RQ3X-ToM是否带来更高水平的合理人类信任与依赖?
- RQ4根据用户感知,哪些图像区域在AI检测特定身体部位时最具影响力?
- RQ5不同解释类型在解释质量(实用性、充分性、细节)方面有何差异?
主要发现
- X-ToM在解释满意度方面显著优于QA和显著性图基线,尤其在实用性、充分性和细节方面(p<0.01)。
- X-ToM与基线方法在响应时间上无显著差异,表明X-ToM解释的处理速度并未更慢。
- 主观依赖度在X-ToM组中高于显著性图基线,表明用户信任度更强。
- 各组在信心度、可理解性、准确性或解释一致性方面均无显著差异。
- 定性依赖度评分与定量依赖度测量结果一致,强化了研究发现的有效性。
- 结果表明,心智理论解释可在不损害可用性或速度的前提下提升信任度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。