[论文解读] Auditing large language models: a three-layered approach
本文提出了一种三层审计框架,用于大型语言模型(LLMs),包括对技术提供商的治理审计、模型发布前的预训练后审计,以及对下游应用的使用审计。该方法能够在技术、伦理和法律维度上实现协调、可行且有效的风险识别,为利益相关方提供一套结构化的工具包,以管理大型语言模型中的新兴风险。
Large language models (LLMs) represent a major advance in artificial intelligence (AI) research. However, the widespread use of LLMs is also coupled with significant ethical and social challenges. Previous research has pointed towards auditing as a promising governance mechanism to help ensure that AI systems are designed and deployed in ways that are ethical, legal, and technically robust. However, existing auditing procedures fail to address the governance challenges posed by LLMs, which display emergent capabilities and are adaptable to a wide range of downstream tasks. In this article, we address that gap by outlining a novel blueprint for how to audit LLMs. Specifically, we propose a three-layered approach, whereby governance audits (of technology providers that design and disseminate LLMs), model audits (of LLMs after pre-training but prior to their release), and application audits (of applications based on LLMs) complement and inform each other. We show how audits, when conducted in a structured and coordinated manner on all three levels, can be a feasible and effective mechanism for identifying and managing some of the ethical and social risks posed by LLMs. However, it is important to remain realistic about what auditing can reasonably be expected to achieve. Therefore, we discuss the limitations not only of our three-layered approach but also of the prospect of auditing LLMs at all. Ultimately, this article seeks to expand the methodological toolkit available to technology providers and policymakers who wish to analyse and evaluate LLMs from technical, ethical, and legal perspectives.
研究动机与目标
- 弥补现有审计实践在应对大型语言模型(LLMs)的涌现特性和通用能力方面存在的不足。
- 开发一种治理机制,通过在模型开发和部署阶段识别风险,补充应用层面的审计。
- 创建一种结构化、协调且可行的审计方法,以应对基础模型和LLMs带来的独特挑战。
- 为技术提供商、政策制定者和审计人员提供一份实用蓝图,从技术、伦理和法律角度评估LLMs。
- 承认并讨论在LLMs背景下审计作为治理工具的局限性,以设定现实预期。
提出的方法
- 在五个数据库(Google Scholar、Scopus、SSRN、Web of Science、arXiv)中系统性地检索文献,关键词涵盖人工智能审计、公平性、透明度和语言模型。
- 采用滚雪球法,从纳入研究的参考文献中识别非学术性审计程序,扩大对学术来源之外的覆盖范围。
- 基于风险/合规性、内部/外部、事前/事后,以及功能/代码/影响等维度,构建审计程序的分类体系。
- 对LLM特有的治理挑战与现有审计能力之间的差距进行分析,提出六项关键设计主张,以实现有效的LLM审计。
- 通过选择在覆盖主要风险、可实施性和成本效益方面联合充分的审计程序,整合出一套最小但高效的审计程序集合。
- 通过三角验证法对草案蓝图进行验证,包括对来自不同司法管辖区的20多个研究人员、审计人员、AI开发者和政策制定者的半结构化访谈。

实验结果
研究问题
- RQ1鉴于大型语言模型具有涌现能力和广泛适应性,如何结构化审计以有效应对其所带来的伦理和社会风险?
- RQ2为确保从模型开发到下游部署的整个LLM生命周期中实现问责和风险缓解,需要哪些审计层级?
- RQ3如何协调治理、模型和应用审计,以在不重复工作的情况下增强透明度、鲁棒性和合规性?
- RQ4哪些实际和制度性约束限制了LLM审计的可行性,又如何在治理框架中加以解决?
- RQ5审计在多大程度上可作为基础模型的可行治理机制?其固有局限性是什么?
主要发现
- 三层架构——治理审计、模型审计和应用审计——提供了一个全面且协调的框架,可在LLM开发和部署的所有阶段识别风险。
- 该框架通过在模型发布前于提供方层面整合审计,实现风险的主动识别,确保伦理和技术防护措施从一开始就嵌入系统。
- 模型审计在预训练完成后、发布前进行,可对安全性、公平性和鲁棒性进行系统性评估,确保在公开部署前达到标准。
- 应用审计确保对LLM下游应用的评估具有情境相关性,捕捉在模型层面评估中可能未显现的风险。
- 该方法设计为可行且可扩展,各审计职能相互重叠但不冗余,优化了资源使用并明确了责任分工。
- 尽管框架具有诸多优势,但仍承认仅靠审计无法完全解决模型不透明性、涌现行为或责任空白等问题,尤其是在模型与环境复杂交互导致损害的情况下。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。