[论文解读] Actionable Guidance for High-Consequence AI Risk Management: Towards Standards Addressing AI Catastrophic Risks
本文通过将NIST人工智能风险管理框架(AI RMF)转化为具体的、可操作的风险管理实践,为管理高后果人工智能风险——尤其是灾难性社会规模危害——提供了切实可行的指导。文章详细提出了关于识别滥用和非预期使用、将灾难性风险因素整合到评估中、减轻人权损害以及报告风险因素的建议,旨在加强人工智能标准并推动关于系统性人工智能安全的对话。
Artificial intelligence (AI) systems can provide many beneficial capabilities but also risks of adverse events. Some AI systems could present risks of events with very high or catastrophic consequences at societal scale. The US National Institute of Standards and Technology (NIST) has been developing the NIST Artificial Intelligence Risk Management Framework (AI RMF) as voluntary guidance on AI risk assessment and management for AI developers and others. For addressing risks of events with catastrophic consequences, NIST indicated a need to translate from high level principles to actionable risk management guidance. In this document, we provide detailed actionable-guidance recommendations focused on identifying and managing risks of events with very high or catastrophic consequences, intended as a risk management practices resource for NIST for AI RMF version 1.0 (released in January 2023), or for AI RMF users, or for other AI risk management guidance and standards as appropriate. We also provide our methodology for our recommendations. We provide actionable-guidance recommendations for AI RMF 1.0 on: identifying risks from potential unintended uses and misuses of AI systems; including catastrophic-risk factors within the scope of risk assessments and impact assessments; identifying and mitigating human rights harms; and reporting information on AI risk factors including catastrophic-risk factors. In addition, we provide recommendations on additional issues for a roadmap for later versions of the AI RMF or supplementary publications. These include: providing an AI RMF Profile with supplementary guidance for cutting-edge increasingly multi-purpose or general-purpose AI. We aim for this work to be a concrete risk-management practices contribution, and to stimulate constructive dialogue on how to address catastrophic risks and associated issues in AI standards.
研究动机与目标
- 弥合高层级人工智能风险原则向具体、可实施的实践转化的空白,以管理灾难性人工智能风险。
- 通过提供针对高后果风险的详细、可操作的建议,支持NIST AI RMF 1.0版本。
- 通过将灾难性风险因素嵌入标准人工智能风险和影响评估中,增强风险评估框架。
- 通过推动在人工智能开发和部署过程中系统性报告灾难性风险因素,加强人工智能治理。
- 通过为未来人工智能标准提出通用型和多用途人工智能系统的发展路线图,为后续AI RMF版本奠定基础。
提出的方法
- 开发了一套结构化方法,将NIST AI RMF原则转化为具体、可操作的风险管理实践。
- 识别了关键风险类别,包括非预期使用、滥用以及先进人工智能系统带来的系统性社会危害。
- 提出了将灾难性风险因素整合到AI RMF框架内现有人工智能风险和影响评估流程中的方案。
- 制定了针对高后果人工智能应用的人权影响评估指导。
- 提出了披露人工智能风险因素(包括具有灾难性潜力的因素)的报告框架。
- 提出了适用于尖端通用型或多用途人工智能系统的AI RMF配置文件,以指导未来标准化工作。
实验结果
研究问题
- RQ1如何将NIST AI RMF中的高层级人工智能风险原则转化为具体、可操作的风险管理实践?
- RQ2在人工智能开发过程中,应系统性识别和管理哪些与灾难性社会后果相关的特定风险因素?
- RQ3如何在风险管理体系内主动评估并减轻人工智能系统的滥用和非预期使用?
- RQ4人权影响评估在管理高后果人工智能风险中应发挥何种作用?
- RQ5人工智能开发者和组织应如何有效报告灾难性风险因素,以支持透明度和问责制?
主要发现
- 本文在NIST AI RMF框架内,为识别和管理灾难性人工智能风险,提供了一套全面的可操作建议。
- 它明确建立了将灾难性风险因素整合到标准人工智能风险和影响评估中的路径,提升了系统性风险意识。
- 作者提出了系统化的风险因素报告方法,包括具有高社会后果的因素,增强了透明度。
- 指导中包含为通用型和多用途人工智能系统提出的AI RMF配置文件,支持未来的标准化工作。
- 该工作通过提供一个实用且基于证据的人工智能风险管理体系,推动了人工智能风险管理的发展,以应对系统性与高影响风险。
- 这些建议与NIST AI RMF 1.0最终版本保持一致,并旨在支持其实施,从而增强其对开发者和监管机构的实际可用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。