[论文解读] Where's the Liability in Harmful AI Speech?
本文探讨了大型语言模型生成有害言论在美国法律下的责任问题,认为现行法律(如第230条)提供的豁免权模糊不清,责任在很大程度上取决于技术设计选择。文章提出一种细致入微、与设计相关的监管方法,通过量身定制的法律标准和最佳实践,激励更安全的人工智能开发。
Generative AI, in particular text-based "foundation models" (large models trained on a huge variety of information including the internet), can generate speech that could be problematic under a wide range of liability regimes. Machine learning practitioners regularly "red team" models to identify and mitigate such problematic speech: from "hallucinations" falsely accusing people of serious misconduct to recipes for constructing an atomic bomb. A key question is whether these red-teamed behaviors actually present any liability risk for model creators and deployers under U.S. law, incentivizing investments in safety mechanisms. We examine three liability regimes, tying them to common examples of red-teamed model behaviors: defamation, speech integral to criminal conduct, and wrongful death. We find that any Section 230 immunity analysis or downstream liability analysis is intimately wrapped up in the technical details of algorithm design. And there are many roadblocks to truly finding models (and their associated parties) liable for generated speech. We argue that AI should not be categorically immune from liability in these scenarios and that as courts grapple with the already fine-grained complexities of platform algorithms, the technical details of generative AI loom above with thornier questions. Courts and policymakers should think carefully about what technical design incentives they create as they evaluate these issues.
研究动机与目标
- 评估生成式人工智能模型在何种情况下以及如何就其有害或虚假言论承担法律责任。
- 研究技术设计决策(如检索增强生成或基于人类反馈的强化学习)如何影响法律责任及第230条的豁免权。
- 识别现行法律框架中未能考虑人工智能模型行为技术细节的漏洞。
- 提出一种精细化的法律方法,使责任标准与负责任的人工智能开发及技术缓解策略相一致。
提出的方法
- 分析三种核心责任制度:诽谤、与犯罪行为直接相关的言论,以及非正常死亡。
- 将特定的红队测试场景(如虚假指控、制造炸弹的说明)映射到法律原则与技术模型行为上。
- 评估第230条豁免权在不同模型架构(包括抽取式、检索增强式及微调模型)中的适用性。
- 评估意图与“犯罪意图”在确定责任中的作用,特别是在虚假事实陈述案件中。
- 提出一套技术最佳实践框架——如免责声明、对持续输出的通知与删除机制,以及特定主题的限制——作为法律上的安全港。
- 建议从广泛豁免转向细粒度的责任制度,以奖励技术上合理、注重安全的设计选择。
实验结果
研究问题
- RQ1基础模型的技术架构在多大程度上影响其在美国法律下的法律责任?
- RQ2第230条豁免权在多大程度上适用于生成式人工智能输出?其适用性如何随不同模型设计而变化?
- RQ3在人工智能缺乏意识的情况下,意图或“心理状态”在认定人工智能生成有害言论责任时起什么作用?
- RQ4检索增强生成或强化学习从人类反馈等技术缓解策略能否构成法律上的安全港?
- RQ5法律责任制度应如何重构,以激励开发更安全、更准确的人工智能系统?
主要发现
- 第230条豁免权并非普遍适用于基础模型,其适用性在很大程度上取决于输出内容是否被视为第三方内容,还是由模型自身生成。
- 采用检索增强生成或抽取式方法的模型比纯生成式模型更可能符合第230条的保护条件。
- 在诽谤、犯罪行为及非正常死亡制度下,对虚假或有害言论的责任是可能的,但取决于技术设计和意图归属。
- 人工智能缺乏意识导致传统‘犯罪意图’要求难以适用,使得责任认定高度依赖系统设计与控制机制。
- 若正式采纳,免责声明、对持续输出的删除机制以及主题特定限制等最佳实践可作为法律上的安全港。
- 现行法律可能无意中激励次优的技术选择——例如为维持豁免权而避免部署安全功能——从而损害公众对更安全人工智能的期待。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。