Skip to main content
QUICK REVIEW

[论文解读] Never trust, always verify : a roadmap for Trustworthy AI?

Lionel Nganyewou Tidjon, Foutse Khomh|arXiv (Cornell University)|Jun 23, 2022
Ethics and Social Impacts of AI被引用 7
一句话总结

本文提出了一种面向人工智能系统的信任与零信任模型,以应对人工智能可信性方面的挑战,强调持续验证、伦理运维(EthicsOps)以及贯穿人工智能生命周期的一套12项核心属性(如透明性、公平性和安全性)。该框架整合了策略执行与自动化验证工具,以确保持续可信,主要贡献包括一个结构化的可信人工智能模型以及一项实用的实施路线图。

ABSTRACT

Artificial Intelligence (AI) is becoming the corner stone of many systems used in our daily lives such as autonomous vehicles, healthcare systems, and unmanned aircraft systems. Machine Learning is a field of AI that enables systems to learn from data and make decisions on new data based on models to achieve a given goal. The stochastic nature of AI models makes verification and validation tasks challenging. Moreover, there are intrinsic biaises in AI models such as reproductibility bias, selection bias (e.g., races, genders, color), and reporting bias (i.e., results that do not reflect the reality). Increasingly, there is also a particular attention to the ethical, legal, and societal impacts of AI. AI systems are difficult to audit and certify because of their black-box nature. They also appear to be vulnerable to threats; AI systems can misbehave when untrusted data are given, making them insecure and unsafe. Governments, national and international organizations have proposed several principles to overcome these challenges but their applications in practice are limited and there are different interpretations in the principles that can bias implementations. In this paper, we examine trust in the context of AI-based systems to understand what it means for an AI system to be trustworthy and identify actions that need to be undertaken to ensure that AI systems are trustworthy. To achieve this goal, we first review existing approaches proposed for ensuring the trustworthiness of AI systems, in order to identify potential conceptual gaps in understanding what trustworthy AI is. Then, we suggest a trust (resp. zero-trust) model for AI and suggest a set of properties that should be satisfied to ensure the trustworthiness of AI systems.

研究动机与目标

  • 分析基于人工智能的系统中的信任概念,并阐明人工智能可信的含义。
  • 识别现有可信人工智能方法中的概念空白,特别是在伦理原则的解释与应用方面。
  • 提出一个全面的信任(及零信任)模型,以确保人工智能生命周期中的端到端可信性。
  • 定义一组12项理想属性(如透明性、公平性和可持续性),这些属性必须满足才能使人工智能系统被视为可信。
  • 整合EthicsOps与基于策略的验证机制,以实现对人工智能可信性的持续监控与验证。

提出的方法

  • 构建了一个包含六个组件的信任与零信任人工智能模型:人类(信任方)、人工智能系统(受信方)、可信人工智能原则(TAI principles)、策略执行点(PEP)、策略决策点(PDP)以及EthicsOps。
  • 整合EthicsOps以实现对人工智能系统行为的持续监控与实时验证,确保在部署过程中持续维持信任。
  • 使用策略引擎,基于可信性检查动态授权对数据和预测的访问,应用TAI原则(如公平性、隐私)进行控制。
  • 采用形式化验证工具(如VeriDeep、DeepZ、RefineZono和RefinePoly)验证模型行为,并检测对抗性攻击下的鲁棒性。
  • 对来自全球组织的100份报告进行人工分析,提取并分类可信人工智能原则,通过关键词提取与冗余控制确保准确性。
  • 在人工智能流水线的每个阶段(包括数据、模型和部署阶段)建立验证层,以维持持续的可信状态。

实验结果

研究问题

  • RQ1当前对可信人工智能的理解中存在哪些概念空白,特别是在伦理原则的解释与实施方面?
  • RQ2如何将零信任模型适配于人工智能系统,以实现持续验证并减少对不可信人工智能决策的依赖?
  • RQ3在不同应用领域中,确保人工智能系统可信性的核心属性有哪些?
  • RQ4如何实现EthicsOps以在人工智能生命周期中持续维护信任,特别是在生产环境中?
  • RQ5策略执行与自动化验证工具在动态和对抗性条件下维持人工智能系统可信性方面发挥什么作用?

主要发现

  • 透明性是在全球100份组织报告中被引用频率最高的原则,表明其在可信人工智能中的核心地位。
  • 本研究识别出12项理想的可信属性:透明性、隐私、公平性、安全性、可靠性、责任、问责制、可解释性、福祉、人权、包容性与可持续性。
  • 所提出的零信任人工智能(ZTA)模型通过基于策略的访问控制与EthicsOps的实时监控,实现了持续验证。
  • 验证工具如VeriDeep、DeepZ与RefineZono在验证模型行为与检测对抗性漏洞方面表现出有效性。
  • 将策略引擎与伦理准则相结合,可实现基于系统状态与输入动态变化的上下文感知信任决策。
  • 对100份报告的两次重复人工分析验证了关键词提取的一致性,并有效减少了冗余,提升了原则分类的可靠性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。