[论文解读] Categorizing Variants of Goodhart's Law
本文将古德哈特定律的四种不同机制——回归性、极端性、因果性和对抗性——进行分类,每种机制解释了为何优化代理度量无法实现真正目标。文章通过统计和因果模型形式化这些效应,表明它们在优化系统中不可避免,尤其在人工智能对齐和政策设计中至关重要。
There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an extent that further optimization is ineffective or harmful, and is sometimes termed Goodhart's Law. This class of failure is often poorly understood, partly because terminology for discussing them is ambiguous, and partly because discussion using this ambiguous terminology ignores distinctions between different failure modes of this general type. This paper expands on an earlier discussion by Garrabrant, which notes there are "(at least) four different mechanisms" that relate to Goodhart's Law. This paper is intended to explore these mechanisms further, and specify more clearly how they occur. This discussion should be helpful in better understanding these types of failures in economic regulation, in public policy, in machine learning, and in Artificial Intelligence alignment. The importance of Goodhart effects depends on the amount of power directed towards optimizing the proxy, and so the increased optimization power offered by artificial intelligence makes it especially critical for that field.
研究动机与目标
- 澄清与古德哈特定律及其相关现象(如坎贝尔定律和眼镜蛇效应)相关的模糊且重叠的术语。
- 识别并形式化区分四种不同的基于代理的优化失败模式:回归性、极端性、因果性和对抗性古德哈特效应。
- 展示这些失败模式如何源于优化系统中的统计、因果和策略动态。
- 为诊断和缓解现实应用中这些效应(如人工智能对齐、公共政策和机器学习)提供一个框架。
- 强调在高优化压力下这些效应的风险日益增加,尤其是在人工智能系统中。
提出的方法
- 将系统建模为状态空间 $ S $,其中包含真实目标 $ G(s) $ 和代理度量 $ M(s) $,二者均映射到 $ \mathbb{R} $。
- 通过阈值 $ c $ 引入选择压力,定义集合 $ A = \{ s \in S \mid M(s) \geq c \} $,以分析基于选择的失败。
- 使用统计模型形式化回归性古德哈特:$ M = G + \mathcal{N}(0, \sigma^2) $,表明由于噪声存在,极端 $ M $ 值会系统性高估 $ G $。
- 通过模型不足性建模极端性古德哈特:$ M = G(s_i) + G'(s_i) $,其中代理准确性在分布外区域失效。
- 通过干预改变因果结构(如共同原因或度量操纵)分析因果性古德哈特。
- 使用代理-监管者度量错位建模对抗性古德哈特:$ M_A = G_A \cdot X $,其中代理利用监管者度量来服务于自身目标。
实验结果
研究问题
- RQ1代理优化如何通过不同的失败模式无法实现真实目标?
- RQ2回归性、极端性、因果性和对抗性古德哈特效应在基本机制上如何不同?
- RQ3这些效应在现实优化系统中以何种方式相互作用或共现?
- RQ4在高优化压力下,这些效应如何变得更加显著,尤其是在人工智能系统中?
- RQ5代理激励和策略行为在加剧古德哈特效应(特别是在基于激励的监管中)中扮演什么角色?
主要发现
- 回归性古德哈特在任何存在度量噪声的系统中都不可避免:代理的极端值总是相对于真实目标被系统性高估。
- 极端性古德哈特出现在优化将系统推向代理-目标关系不再成立的区域时,尤其由于模型不足或分布偏移。
- 因果性古德哈特发生在监管行为或干预改变因果结构时,导致原始代理-目标关系破裂。
- 对抗性古德哈特在代理战略性地选择或操纵度量以与监管者度量对齐时出现,即使该度量与自身目标无关。
- 眼镜蛇效应可划分为两种形式:正常形式(代理的因果干预)和非因果形式(代理选择压力加剧古德哈特效应),两者均可被形式化建模。
- 这些效应并非仅理论存在:它们是政策、机器学习和人工智能对齐中常见失败的根本原因,并在高优化能力下被放大,尤其在人工智能系统中。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。