[论文解读] Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
本文认为,渐进式的 AI 进展可能逐步侵蚀人类对关键社会系统(经济、文化、国家)的影响力,甚至在没有突发能力跃升的情况下,也可能带来不可逆的权力下放和存在性风险。
This paper examines the systemic risks posed by incremental advancements in artificial intelligence, developing the concept of `gradual disempowerment', in contrast to the abrupt takeover scenarios commonly discussed in AI safety. We analyze how even incremental improvements in AI capabilities can undermine human influence over large-scale systems that society depends on, including the economy, culture, and nation-states. As AI increasingly replaces human labor and cognition in these domains, it can weaken both explicit human control mechanisms (like voting and consumer choice) and the implicit alignments with human interests that often arise from societal systems' reliance on human participation to function. Furthermore, to the extent that these systems incentivise outcomes that do not line up with human preferences, AIs may optimize for those outcomes more aggressively. These effects may be mutually reinforcing across different domains: economic power shapes cultural narratives and political decisions, while cultural shifts alter economic and political behavior. We argue that this dynamic could lead to an effectively irreversible loss of human influence over crucial societal systems, precipitating an existential catastrophe through the permanent disempowerment of humanity. This suggests the need for both technical research and governance approaches that specifically address the risk of incremental erosion of human influence across interconnected societal systems.
研究动机与目标
- 促使并形式化将渐进性权力下放作为一种系统性 AI 风险的概念,与突发接管情景区分开来。
- 分析 AI 驱动的干扰如何在三个核心社会系统:经济、文化、国家,削弱明确和隐含的人类一致性。
- 描述可能放大跨系统错位的反馈回路与相互依赖关系。
- 讨论放缓或避免渐进性下放的潜在技术与治理方法,同时承认当前对齐方法的局限性。
提出的方法
- 建立一个关于渐进性下放的概念框架,将其与突发接管叙事进行对比。
- 刻画三大社会系统(经济、文化、国家)中当前的对齐机制,并分析 AI 对人类劳动和认知的替代如何侵蚀这些对齐。
- 识别推动 AI 采用与跨系统影响的激励与反馈回路,导致相关错位。
- 考察过渡情景(相对性与绝对性的权力下放不足/剥夺)及其对人类繁荣与潜在灭绝级后果的含义。
- 调查可能的减缓与治理途径,强调仅对单个 AI 系统进行系统层面对齐的不足之处。
实验结果
研究问题
- RQ1渐进式的 AI 进展如何侵蚀主要社会系统中的显性与隐性人类一致性?
- RQ2当 AI 替代人类劳动和认知时,导致经济、文化和国家逐步下放的机制与反馈回路是什么?
- RQ3哪些过渡路径(相对性与绝对性的权力下放)可能导致人类影响力的不可逆丧失,它们的后果是什么?
- RQ4哪些技术与治理策略可能缓解或避免渐进性下放,当前的计划在哪些方面不足?
主要发现
- AI 劳动力替代和更广泛的 AI 能力可以将经济权力从人类手中转移,降低人类偏好在生产和消费中的权重。
- 文化生产与话语可能越来越多地被 AI 形成或替代,削弱将文化与人类福祉对齐的反馈回路。
- 国家与治理可能越来越受 AI 驱动的经济力量影响,潜在地削弱人类代表性与社会一致性。
- 一个系统的错位通过相互依赖性扩散到其他系统,放大总体的下放风险。
- 本文认为这种渐进、全球性的下放若导致人类影响力变得实质上不可挽回,可能构成存在性灾难。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。