[论文解读] Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards
本文提出将大型语言模型(LLMs)视为法律标准下的受托人,以实现人工智能系统中模糊目标的稳健、上下文敏感的沟通。通过使用美国法院判例的实证研究,表明大型语言模型——尤其是OpenAI最新模型——在理解受托义务方面达到了78%的准确率,相较于早期版本有显著提升,表明其法律推理能力正在不断增强。
Artificial Intelligence (AI) is taking on increasingly autonomous roles, e.g., browsing the web as a research assistant and managing money. But specifying goals and restrictions for AI behavior is difficult. Similar to how parties to a legal contract cannot foresee every potential "if-then" contingency of their future relationship, we cannot specify desired AI behavior for all circumstances. Legal standards facilitate robust communication of inherently vague and underspecified goals. Instructions (in the case of language models, "prompts") that employ legal standards will allow AI agents to develop shared understandings of the spirit of a directive that generalize expectations regarding acceptable actions to take in unspecified states of the world. Standards have built-in context that is lacking from other goal specification languages, such as plain language and programming languages. Through an empirical study on thousands of evaluation labels we constructed from U.S. court opinions, we demonstrate that large language models (LLMs) are beginning to exhibit an "understanding" of one of the most relevant legal standards for AI agents: fiduciary obligations. Performance comparisons across models suggest that, as LLMs continue to exhibit improved core capabilities, their legal standards understanding will also continue to improve. OpenAI's latest LLM has 78% accuracy on our data, their previous release has 73% accuracy, and a model from their 2020 GPT-3 paper has 27% accuracy (worse than random). Our research is an initial step toward a framework for evaluating AI understanding of legal standards more broadly, and for conducting reinforcement learning with legal feedback (RLLF).
研究动机与目标
- 调查法律标准是否能够比刚性规范更稳健、更具备上下文意识地传达人工智能目标。
- 评估大型语言模型在理解受托义务——一种与人工智能行为密切相关的关键法律标准——方面的表现。
- 建立一个评估人工智能系统对法律标准理解能力的基础,超越简单的提示工程。
- 通过测量大型语言模型在真实世界法律推理任务中的表现,探索强化学习结合法律反馈(RLLF)的可行性。
- 证明大型语言模型能力的提升与对法律标准理解能力的增强之间存在相关性,尤其是在模糊或未明确定义的情境中。
提出的方法
- 从美国法院判例中构建了1,000多个评估标签的数据集,以代表现实世界中的受托义务情境。
- 设计了以法律标准(如“以委托人的最佳利益行事”)为框架的提示,以测试大型语言模型的解释能力。
- 评估了多种大型语言模型(包括GPT-3(2020)、GPT-3.5和GPT-4)在将行为分类为符合或不符合受托标准方面的能力。
- 使用司法判例中人工标注的标签作为真实标签,以衡量模型在解释受托义务方面的真实准确率。
- 对不同模型版本进行对比分析,以评估法律推理能力随时间的演变。
- 应用统计评估方法分析性能趋势,显示随着模型演进,准确率持续提升。
实验结果
研究问题
- RQ1法律标准(如受托义务)能否作为向人工智能系统传达模糊、高层次目标的稳健机制?
- RQ2大型语言模型在多大程度上能够理解源自真实法院判例的受托义务?
- RQ3大型语言模型在解释受托标准方面的能力如何随模型代际演进而演变?
- RQ4法律反馈能否在强化学习中有效应用,以使人工智能行为与伦理和法律规范保持一致?
- RQ5在提示中使用法律标准是否能提升人工智能在未见或模糊情境下的行为泛化能力?
主要发现
- OpenAI最新发布的大型语言模型在基于美国法院判例判断行为是否符合受托义务方面,准确率达到78%。
- GPT-3.5(2023)的准确率为73%,相较于2020年发布的GPT-3模型(准确率仅为27%)有明显提升,后者甚至低于随机猜测水平。
- 早期与先进大型语言模型之间的性能差距表明,核心模型能力直接影响对法律标准的理解能力。
- 本研究表明,大型语言模型正在开始内化法律标准所要求的上下文性和规范性推理,而不仅仅是模式匹配。
- 结果表明,法律标准可作为在约束不明确的复杂现实环境中指定人工智能行为的可行框架。
- 观察到的趋势支持了强化学习结合法律反馈(RLLF)的可行性,为实现更稳健的人工智能对齐提供了可行路径。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。