[论文解读] System Cards for AI-Based Decision-Making for Public Policy
本文提出了一套系统问责基准和系统卡片,以增强基于人工智能的自动化决策系统(ADSs)在透明度和问责性方面的表现。该框架采用56项标准,按数据、模型、代码和系统四个维度组织成四乘四矩阵,列分别为开发、评估、缓解和保障,支持结构化审计,并生成标准化的系统卡片,适用于公共政策应用。
Decisions impacting human lives are increasingly being made or assisted by automated decision-making algorithms. Many of these algorithms process personal data for predicting recidivism, credit risk analysis, identifying individuals using face recognition, and more. While potentially improving efficiency and effectiveness, such algorithms are not inherently free from bias, opaqueness, lack of explainability, maleficence, and the like. Given that the outcomes of these algorithms have a significant impact on individuals and society and are open to analysis and contestation after deployment, such issues must be accounted for before deployment. Formal audits are a way of ensuring algorithms meet the appropriate accountability standards. This work, based on an extensive analysis of the literature and an expert focus group study, proposes a unifying framework for a system accountability benchmark for formal audits of artificial intelligence-based decision-aiding systems. This work also proposes system cards to serve as scorecards presenting the outcomes of such audits. It consists of 56 criteria organized within a four-by-four matrix composed of rows focused on (i) data, (ii) model, (iii) code, (iv) system, and columns focused on (a) development, (b) assessment, (c) mitigation, and (d) assurance. The proposed system accountability benchmark reflects the state-of-the-art developments for accountable systems, serves as a checklist for algorithm audits, and paves the way for sequential work in future research.
研究动机与目标
- 为应对在高风险公共政策领域中使用基于人工智能的自动化决策系统(ADSs)所引发的偏见、不透明性和可解释性不足等问题。
- 开发一种标准化、可审计的框架,用于在部署前评估人工智能系统的问责性。
- 创建一种实用工具——系统卡片,将审计结果提炼为对利益相关者可访问、结构化的报告。
- 通过嵌入问责原则,支持人工智能系统从设计到部署的全生命周期管理。
- 通过提供具体且可操作的基准,推动算法问责领域的发展,为研究人员和从业者提供支持。
提出的方法
- 提出一个由56项标准构成的系统问责基准,组织为四乘四矩阵:行代表数据、模型、代码和系统,列代表开发、评估、缓解和保障。
- 设计系统卡片作为标准化的输出成果,使用基准标准总结审计结果。
- 结合文献综述和专家焦点小组的见解,确保与最先进的问责实践保持一致。
- 结构化标准以支持定性和定量评估,承认在人工智能保险或团队多样性等新兴概念上测量存在局限。
- 强调审计结果必须在社会和法律语境中解读,避免采用简单化的通过/失败或单一评分评估。
- 将该框架定位为可动态演进的工具,支持未来指南的制定和自动化审计工具的开发。
实验结果
研究问题
- RQ1如何设计一个全面且标准化的框架,以审计公共政策中的人工智能自动化决策系统?
- RQ2在数据、模型、代码和系统各层中,确保问责性、公平性和透明性的关键标准是什么?
- RQ3如何通过标准化且可解释的格式(如系统卡片)有效向利益相关者传达审计结果?
- RQ4在涉及多种公平性概念或定性评估时,测量问责标准的实际挑战是什么?
- RQ5该框架如何适应不同类型的人工智能系统,包括使用无监督学习或基于规则逻辑的系统?
主要发现
- 系统问责基准包含56项标准,组织为四乘四矩阵,支持对人工智能系统在数据、模型、代码和系统维度的系统性评估。
- 提出系统卡片作为标准化、可审计的输出成果,使用基准标准总结审计结果,提升透明度和利益相关者理解。
- 该框架并非旨在成为“一刀切”的解决方案,而是一个持续演进的工具,用于指导人工智能系统全生命周期中的问责实践。
- 基准承认,并非所有标准都能客观测量——部分标准需要定性评估,特别是文档工作和人工智能保险等新兴概念。
- 该框架强调,公平性评估具有情境依赖性,由于公平性概念之间存在冲突且涉及价值判断,不存在单一客观方法。
- 研究发现,复杂且不透明的系统——尤其是使用深度学习或历史数据的系统——在可审计性方面面临重大挑战,即使未显式包含人口统计特征,也可能重现系统性偏见。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。