Skip to main content
QUICK REVIEW

[论文解读] Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI

Alon Jacovi, Ana Marasović|arXiv (Cornell University)|Oct 15, 2020
Explainable Artificial Intelligence (XAI)参考文献 71被引用 66
一句话总结

本论文将受人际信任启发的人工智能与人类之间的信任形式化,提出契约信任和可信性,并区分有理信任与不当信任,将信任与 XAI 与评估联系起来。

ABSTRACT

Trust is a central component of the interaction between people and AI, in that 'incorrect' levels of trust may cause misuse, abuse or disuse of the technology. But what, precisely, is the nature of trust in AI? What are the prerequisites and goals of the cognitive mechanism of trust, and how can we promote them, or assess whether they are being satisfied in a given interaction? This work aims to answer these questions. We discuss a model of trust inspired by, but not identical to, sociology's interpersonal trust (i.e., trust between people). This model rests on two key properties of the vulnerability of the user and the ability to anticipate the impact of the AI model's decisions. We incorporate a formalization of 'contractual trust', such that trust between a user and an AI is trust that some implicit or explicit contract will hold, and a formalization of 'trustworthiness' (which detaches from the notion of trustworthiness in sociology), and with it concepts of 'warranted' and 'unwarranted' trust. We then present the possible causes of warranted trust as intrinsic reasoning and extrinsic behavior, and discuss how to design trustworthy AI, how to evaluate whether trust has manifested, and whether it is warranted. Finally, we elucidate the connection between trust and XAI using our formalization.

研究动机与目标

  • 将受社会学人际信任启发的人与人工智能模型之间的信任进行定义。
  • 引入契约信任并区分可信性。
  • 区分有理信任与不当信任并讨论其含义。
  • 解释信任如何与 XAI 与模型评估相关。
  • 提出在现实世界互动中设计和评估可信AI的框架。

提出的方法

  • 采用信任的两个核心属性:用户脆弱性以及预测AI影响的能力。
  • 将契约信任形式化,以明确AI被信任去做的事情。
  • 将可信性与信任区分,并定义有理信任与不当信任。
  • 将内在信任(与内部推理对齐)与外在信任(可置信的外部行为)表征为促进信任的机制。
  • 将该框架与欧洲指南及标准化文档(例如数据表、模型卡)相关联以明确契约。
  • 讨论评估方法(代理解释、部署后数据、测试集)以证明外在信任。

实验结果

研究问题

  • RQ1人工-AI 信任的前提条件是什么?
  • RQ2如何对人机互动中的契约信任和可信性进行形式化定义?
  • RQ3AI 中有理信任与不当信任有何区别?
  • RQ4在实践中如何促进与评估信任,包括 XAI 的作用?
  • RQ5提出的形式化如何与现有指南和文档实践相连?

主要发现

  • 人类与 AI 之间的信任是一个有方向性的交易,要求脆弱性和对影响的预期。
  • 契约信任和可信性可以形式化,以区分何时信任是有理的,何时是无理的。
  • 信任可能依赖于情境,契约取决于互动情境。
  • 信任机制包括内在信任(与用户先验一致的可解释推理)以及外在信任(对评估方法和数据的信任)。
  • 评估方案(代理判断、部署后数据和测试集)对于建立外在信任、评估契约是否能维持至关重要。
  • AI 的可解释性应与用户的先验和契约特定目标保持一致,以培养真实信任,而非仅仅感知。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。