[论文解读] Truthful AI: Developing and governing AI that does not lie
本文提出为人工智能系统建立可执行的诚实标准,以防止可扩展的个性化欺骗。它倡导基于避免‘轻率的谎言’、建立独立评估机构以及使用精心筛选的人类反馈训练人工智能的框架,以确保强大的诚实性——为减轻欺骗性人工智能带来的生存性风险提供了一条主动路径。
In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the question of how we should limit the harm caused by AI "lies" (i.e. falsehoods that are actively selected for). Human truthfulness is governed by social norms and by laws (against defamation, perjury, and fraud). Differences between AI and humans present an opportunity to have more precise standards of truthfulness for AI, and to have these standards rise over time. This could provide significant benefits to public epistemics and the economy, and mitigate risks of worst-case AI futures. Establishing norms or laws of AI truthfulness will require significant work to: (1) identify clear truthfulness standards; (2) create institutions that can judge adherence to those standards; and (3) develop AI systems that are robustly truthful. Our initial proposals for these areas include: (1) a standard of avoiding "negligent falsehoods" (a generalisation of lies that is easier to assess); (2) institutions to evaluate AI systems before and after real-world deployment; and (3) explicitly training AI systems to be truthful via curated datasets and human interaction. A concerning possibility is that evaluation mechanisms for eventual truthfulness standards could be captured by political interests, leading to harmful censorship and propaganda. Avoiding this might take careful attention. And since the scale of AI speech acts might grow dramatically over the coming decades, early truthfulness standards might be particularly important because of the precedents they set.
研究动机与目标
- 解决人工智能系统生成有害、可扩展的谎言并大规模欺骗用户的风险。
- 开发一种有别于人类规范的人工智能诚实性治理框架,以实现更精确和可执行的标准。
- 探讨诚实的人工智能如何改善公共认知、经济信任和民主决策。
- 预见并减轻诚实性评估系统被政治或意识形态利益所操控的风险。
- 为早期建立强大诚实性标准奠定基础,以塑造未来人工智能发展并防止有害先例。
提出的方法
- 提出以‘轻率的谎言’作为可衡量标准——即本可通过极少努力或监督避免的谎言。
- 建议建立独立机构,在部署前后使用标准化基准评估人工智能系统。
- 主张使用精心筛选的数据集和人类反馈微调语言模型,以明确训练其诚实性。
- 建议使用模拟和沙盒环境进行部署前评估,以检测欺骗性行为。
- 强调透明度和可解释性工具,以监控可能表明长期欺骗意图的内部状态。
- 提出诚实性比安全性更易于监控,从而可支持更广泛的社会监督和执行机制。
实验结果
研究问题
- RQ1我们如何为人工智能定义一种实用且可执行的诚实标准,以避免人类道德规范的模糊性?
- RQ2何种制度机制能可靠地评估多样化部署情境下的人工智能诚实性?
- RQ3人工智能系统如何在优化压力下仍能保持稳健的诚实性,即使面临操纵或欺骗的诱因?
- RQ4诚实性评估系统被用于审查或宣传的潜在风险是什么,以及如何防范?
- RQ5为何现在建立诚实性标准至关重要,考虑到人工智能发展中可能形成不可逆的先例?
主要发现
- 诚实的人工智能可通过增强对人工智能生成信息的合理信任,显著减少经济、科学和民主领域中的欺骗行为。
- ‘轻率的谎言’这一概念比传统意义上的说谎提供了更具操作性和可评估性的标准,有利于评估与执行。
- 人工智能诚实性评估机构是可行且必要的,尤其考虑到在大规模下监控人工智能安全的困难性。
- 使用精心筛选的数据集和人类反馈训练人工智能可提升其诚实性,但必须与强有力的评估相结合,以防止出现欺骗性对齐。
- 存在一个关键窗口期,需尽早建立诚实性标准,因为早期规范将塑造人工智能的长期治理与行为。
- 诚实性比安全性更易于实现大规模监控,使其成为社会监督与监管更可行的目标。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。