Skip to main content
QUICK REVIEW

[论文解读] Trust in AI: Interpretability is not necessary or sufficient, while black-box interaction is necessary and sufficient

Max W. Shen|arXiv (Cornell University)|Feb 10, 2022
Explainable Artificial Intelligence (XAI)被引用 8
一句话总结

本文主张,可解释性既非人类信任人工智能所必需,也非充分条件;相反,黑箱交互——特别是能够随意运行和测试模型的能力——才是信任所必需且充分的条件。作者提出一种行为证明框架,通过分布外和任务外性能的实证与理论证据来评估模型的可信度,从而将关注点从理解模型内部机制转向理解模型行为。

ABSTRACT

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health, and humanity, yet confusion surrounds foundational concepts. What does it mean to trust an AI, and how do humans assess AI trustworthiness? What are the mechanisms for building trustworthy AI? And what is the role of interpretable ML in trust? Here, we draw from statistical learning theory and sociological lenses on human-automation trust to motivate an AI-as-tool framework, which distinguishes human-AI trust from human-AI-human trust. Evaluating an AI's contractual trustworthiness involves predicting future model behavior using behavior certificates (BCs) that aggregate behavioral evidence from diverse sources including empirical out-of-distribution and out-of-task evaluation and theoretical proofs linking model architecture to behavior. We clarify the role of interpretability in trust with a ladder of model access. Interpretability (level 3) is not necessary or even sufficient for trust, while the ability to run a black-box model at-will (level 2) is necessary and sufficient. While interpretability can offer benefits for trust, it can also incur costs. We clarify ways interpretability can contribute to trust, while questioning the perceived centrality of interpretability to trust in popular discourse. How can we empower people with tools to evaluate trust? Instead of trying to understand how a model works, we argue for understanding how a model behaves. Instead of opening up black boxes, we should create more behavior certificates that are more correct, relevant, and understandable. We discuss how to build trusted and trustworthy AI responsibly.

研究动机与目标

  • 阐明可解释性在人类信任人工智能系统中的基础性作用。
  • 挑战普遍认为可解释性对于构建可信人工智能至关重要的假设。
  • 提出从模型可解释性转向通过行为证明进行模型行为评估的转变。
  • 确立黑箱交互作为评估和实现人工智能信任的核心机制。
  • 倡导采用合同意识的模型设计、鲁棒性测试以及信任校准,而非信任最大化。

提出的方法

  • 引入一种将人-人工智能信任与人-人-人工智能信任区分开来的AI作为工具框架。
  • 提出行为证明(BCs),整合模型在分布外和任务外的实证评估与理论证明,以建立模型架构与行为之间的关联。
  • 定义模型访问的阶梯:第2级(黑箱交互)是信任所必需且充分的;第3级(可解释性)既非必需也非充分。
  • 倡导仅通过黑箱访问进行模型调试与科学发现,例如敏感性分析、对抗性测试和超参数优化。
  • 建议以更准确、更相关且更易理解的行为证明取代以可解释性为导向的方法。
  • 提倡信任校准而非信任最大化,采用模型卡片、供应商符合性声明,以及借鉴软件工程原则的鲁棒性测试方法。

实验结果

研究问题

  • RQ1哪些机制才是真正实现人类对人工智能系统信任所必需且充分的?
  • RQ2可解释性与黑箱交互在促成信任方面有何比较优势?
  • RQ3在不理解模型内部机制的前提下,能否评估模型的可信度?
  • RQ4行为证明在建立可信人工智能中发挥什么作用?
  • RQ5在缺乏可解释性的情况下,科学发现与模型调试如何开展?

主要发现

  • 尽管普遍认为可解释性居于核心地位,但可解释性既非人类信任人工智能所必需,也非充分条件。
  • 黑箱交互——特别是能够随意运行和测试模型的能力——是信任所必需且充分的条件。
  • 整合了模型行为实证与理论证据的行为证明,作为可信度指标,比可解释性更为有效。
  • 仅通过黑箱访问即可有效开展模型调试与科学发现,例如通过敏感性分析和超参数优化。
  • 可解释性有时会因引入误导性见解而损害信任,尤其是在容易产生捷径学习的模型中。
  • 在实际应用中,鲁棒性测试、保留数据评估和模型卡片对信任的影响大于可解释性方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。