Skip to main content
QUICK REVIEW

[论文解读] The Response Shift Paradigm to Quantify Human Trust in AI Recommendations

Ali Shafti, Victoria Derks|arXiv (Cornell University)|Feb 16, 2022
Explainable Artificial Intelligence (XAI)被引用 5
一句话总结

本文提出了一种定量的人机协同范式,通过追踪'响应变化'——即用户在查看AI推荐和解释后,其初始决策与最终决策之间的变化——来衡量人类对AI推荐的信任度。该方法实现了对可解释AI系统的客观、连续且可扩展的评估,揭示了解释质量显著影响信任度,且人类决策的变化可被量化,从而通过机器学习优化AI的可信度。

ABSTRACT

Explainability, interpretability and how much they affect human trust in AI systems are ultimately problems of human cognition as much as machine learning, yet the effectiveness of AI recommendations and the trust afforded by end-users are typically not evaluated quantitatively. We developed and validated a general purpose Human-AI interaction paradigm which quantifies the impact of AI recommendations on human decisions. In our paradigm we confronted human users with quantitative prediction tasks: asking them for a first response, before confronting them with an AI's recommendations (and explanation), and then asking the human user to provide an updated final response. The difference between final and first responses constitutes the shift or sway in the human decision which we use as metric of the AI's recommendation impact on the human, representing the trust they place on the AI. We evaluated this paradigm on hundreds of users through Amazon Mechanical Turk using a multi-branched experiment confronting users with good/poor AI systems that had good, poor or no explainability. Our proof-of-principle paradigm allows one to quantitatively compare the rapidly growing set of XAI/IAI approaches in terms of their effect on the end-user and opens up the possibility of (machine) learning trust.

研究动机与目标

  • 开发一种定量、客观且可扩展的方法,用于衡量人类对AI推荐的信任度。
  • 评估不同水平的AI可解释性与系统质量对人类决策变化的影响。
  • 构建一个反馈回路,通过量化AI推荐对人类选择的影响,实现信任的机器学习优化。
  • 摒弃主观自评方式,采用连续的行为基信任度指标,用于人机交互中的信任评估。
  • 提供一种通用协议,适用于监督回归任务,用于评估XAI/IAI系统。

提出的方法

  • 该范式采用三阶段人机协同协议:首先,用户提交初始预测;其次,用户查看AI的推荐与解释;第三,用户修改其预测。
  • 关键指标为'响应变化'——即初始响应与最终响应之间的连续差异,用于量化AI对人类决策的影响。
  • 该方法基于贝叶斯决策理论,将人类决策建模为基于AI提供新信息的 probabilistic(概率性)更新。
  • 实验在Amazon Mechanical Turk上进行,涉及数百名用户,使用表格数据预测学生分数。
  • 设计中控制了AI质量(良好 vs. 劣质)和可解释性(良好、劣质或无解释),以实现系统性比较。
  • 响应变化被视为可测量的连续信号,可用于训练或优化AI系统以提升人类信任度。

实验结果

研究问题

  • RQ1AI解释质量在多大程度上影响人类决策变化的幅度,以响应AI推荐?
  • RQ2AI系统性能(良好 vs. 劣质)在多大程度上通过响应变化影响人类信任?
  • RQ3是否能从人机交互中可靠地提取出连续的行为基信任度指标,而无需依赖自评?
  • RQ4对AI系统的训练经验在多大程度上调节响应变化与信任形成?
  • RQ5响应变化信号是否可用于实现AI解释的基于梯度的优化,以提升人类信任?

主要发现

  • 响应变化指标成功量化了AI推荐对人类决策的影响,提供了连续且客观的信任度度量。
  • 解释质量显著影响响应变化,高质量解释导致比低质量或无解释更大的且更一致的决策变化。
  • 当AI准确时,用户更可能将其决策转向AI推荐,即使解释较弱,表明性能在信任形成中比解释质量更具决定性。
  • 该方法揭示了用户对不准确AI系统产生错误信任的案例,尤其是在解释格式良好但内容错误时,凸显了表面信任的风险。
  • 响应变化信号表现出足够的变异性与敏感性,能够检测出不同AI系统质量与解释类型之间的差异,验证了其作为可靠评估指标的有效性。
  • 该框架可反馈至机器学习流程中,使基于可测量的人类响应变化优化AI解释成为可能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。