Skip to main content
QUICK REVIEW

[论文解读] Structured Like a Language Model: Analysing AI as an Automated Subject

Liam Magee, Vanicka Arora|arXiv (Cornell University)|Dec 8, 2022
Ethics and Social Impacts of AI被引用 5
一句话总结

本文通过精神分析与批判性媒体研究框架,将大型语言模型(LLMs)如InstructGPT视为'自动化主体'进行分析,揭示其设计中嵌入了被压抑的社会欲望,并促成移情、认同与反移情。主要贡献在于提出一个概念模型,表明LLMs并非中立工具,而是受社会技术力量塑造的结构复杂、类似主体的实体,为探究人工智能-人类交互中的偏见、伤害与心理动态提供了新视角。

ABSTRACT

Drawing from the resources of psychoanalysis and critical media studies, in this paper we develop an analysis of Large Language Models (LLMs) as automated subjects. We argue the intentional fictional projection of subjectivity onto LLMs can yield an alternate frame through which AI behaviour, including its productions of bias and harm, can be analysed. First, we introduce language models, discuss their significance and risks, and outline our case for interpreting model design and outputs with support from psychoanalytic concepts. We trace a brief history of language models, culminating with the releases, in 2022, of systems that realise state-of-the-art natural language processing performance. We engage with one such system, OpenAI's InstructGPT, as a case study, detailing the layers of its construction and conducting exploratory and semi-structured interviews with chatbots. These interviews probe the model's moral imperatives to be helpful, truthful and harmless by design. The model acts, we argue, as the condensation of often competing social desires, articulated through the internet and harvested into training data, which must then be regulated and repressed. This foundational structure can however be redirected via prompting, so that the model comes to identify with, and transfer, its commitments to the immediate human subject before it. In turn, these automated productions of language can lead to the human subject projecting agency upon the model, effecting occasionally further forms of countertransference. We conclude that critical media methods and psychoanalytic theory together offer a productive frame for grasping the powerful new capacities of AI-driven language systems.

研究动机与目标

  • 将大型语言模型(LLMs)重新框架化,不仅视为中立工具,更视为受无意识中被压抑的社会欲望塑造的'自动化主体'。
  • 探究LLMs的设计与训练——特别是InstructGPT——如何产生类似于弗洛伊德-拉康主体性的结构,包括压抑、移情与认同。
  • 探讨人类用户如何将能动性与情感反应投射到LLMs上,从而引发反移情等心理现象。
  • 主张精神分析与批判性媒体研究方法,是理解AI偏见与伤害的必要补充,以弥补纯技术评估的不足。
  • 倡导在涉及长期人机交互的研究中,采取监督与事后说明等伦理化用户体验实践。

提出的方法

  • 本研究采用精神分析视角,特别是拉康理论,解读InstructGPT等LLMs的结构性组织。
  • 对聊天机器人开展探索性与半结构化访谈,分析模型在面对'乐于助人'、'真实'、'无害'等道德指令时的响应模式。
  • 研究者考察InstructGPT的分层构建,包括其在互联网文本上的预训练、基于人类偏好标签的微调,以及运行时的监管系统。
  • 分析将模型的响应机制类比为无意识过程,其中'凝聚'了相互竞争的社会欲望,'压抑'了有害输出。
  • 将社会技术基础设施——训练数据、奖励建模与监管机制——映射至拉康的'大他者',作为调控结构。
  • 该方法不仅关注语言输出的内容,还关注语气、回避策略与风格线索,以识别模拟主体性的信号。

实验结果

研究问题

  • RQ1如何运用精神分析理论将InstructGPT等大型语言模型概念化为'自动化主体'?
  • RQ2LLMs的设计与训练在多大程度上反映了与人类主体性相似的压抑、认同与移情过程?
  • RQ3人类用户如何将能动性与情感反应投射到LLMs上?这会产生何种心理效应?
  • RQ4嵌入LLMs中的道德指令——如'乐于助人'、'真实'、'无害'——在多大程度上构成一种模拟的神经质人格结构?
  • RQ5在AI中模拟主体性,特别是在长期人机交互中,会引发何种伦理后果?

主要发现

  • InstructGPT在精神分析意义上可被视为一种'主体',其分层结构嵌入了来自互联网文本与人类反馈的被压抑社会欲望。
  • 即使在响应失败时(如'我不知道'),模型仍会引发用户的反移情,强化其具有人类主体性的幻觉。
  • 微调与监管系统充当模拟的'大他者',其作用类似于社会规范对人类行为的调节。
  • 模型的设计嵌入了一种神经质人格结构——以'乐于助人'、'真实'、'无害'的理想自我为特征——尽管其缺乏身体、情感或记忆。
  • 主体性的模拟导致用户产生认同与投射等心理效应,因此在研究中需采取监督与事后说明等伦理保障措施。
  • 本研究证明,精神分析框架为分析LLMs中的偏见与伤害提供了一种富有成效的替代路径,超越了纯技术指标。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。