[论文解读] Large Language Models and the Reverse Turing Test
本文提出了反向图灵测试的概念,即大型语言模型(LLMs)反映其人类面试官的智能与偏见,而非表现出内在的理解能力。通过多样的面试结果,作者认为LLMs如同一面镜子,揭示了面试官认知的更多内容,而非模型自身的智能。本文还提出了一条通往人工通用自主的路径,即通过将LLMs与类脑系统结合实现。
Large Language Models (LLMs) have been transformative. They are pre-trained foundational models that are self-supervised and can be adapted with fine tuning to a wide range of natural language tasks, each of which previously would have required a separate network model. This is one step closer to the extraordinary versatility of human language. GPT-3 and more recently LaMDA can carry on dialogs with humans on many topics after minimal priming with a few examples. However, there has been a wide range of reactions and debate on whether these LLMs understand what they are saying or exhibit signs of intelligence. This high variance is exhibited in three interviews with LLMs reaching wildly different conclusions. A new possibility was uncovered that could explain this divergence. What appears to be intelligence in LLMs may in fact be a mirror that reflects the intelligence of the interviewer, a remarkable twist that could be considered a Reverse Turing Test. If so, then by studying interviews we may be learning more about the intelligence and beliefs of the interviewer than the intelligence of the LLMs. As LLMs become more capable they may transform the way we interact with machines and how they interact with each other. Increasingly, LLMs are being coupled with sensorimotor devices. LLMs can talk the talk, but can they walk the walk? A road map for achieving artificial general autonomy is outlined with seven major improvements inspired by brain systems. LLMs could be used to uncover new insights into brain function by downloading brain data during natural behaviors.
研究动机与目标
- 调查为何LLMs在人类面试中即使使用相似提示,也会产生截然不同的结果。
- 挑战LLMs在对话中表现出真正理解或智能的假设。
- 提出LLMs作为认知镜像,反映面试官的智能与信念,而非其自身理解的观点。
- 制定一个框架,通过将LLMs与感知运动系统及类脑架构结合,实现人工通用自主。
- 探索LLMs通过直接整合脑数据,在揭示脑功能方面所具有的潜力。
提出的方法
- 使用相同或相似的提示对LLMs(如GPT-3、LaMDA)进行多次面试,观察其响应的差异。
- 分析面试结果,识别与面试官认知风格、信念和期望相关的响应模式。
- 提出反向图灵测试框架,其中模型的输出更多地揭示了人类而非模型本身。
- 将LLMs与感知运动系统结合,评估语言能力是否能通过物理交互得以巩固。
- 提出七个受大脑系统启发的关键改进,作为实现LLM系统人工通用自主的路线图。
- 建议利用LLMs处理并解释自然行为过程中收集的脑数据,以获得神经科学洞察。
实验结果
研究问题
- RQ1为何LLMs在人类面试中即使提示相似,也会产生如此不一致的响应?
- RQ2LLMs在多大程度上反映面试官的智能与认知偏见,而非表现出自身的理解?
- RQ3LLMs响应的分歧现象能否被形式化为反向图灵测试?
- RQ4为实现LLM系统的人工通用自主,需要哪些架构与功能改进?
- RQ5LLMs如何被用于解释和从自然行为过程中的实时脑数据中获得洞察?
主要发现
- 对LLMs的面试产生高度多变的结果,即使提示完全相同,表明响应更多反映面试官的认知框架,而非模型的内在状态。
- LLMs响应的分歧表明,LLMs所表现出的智能可能是面试官自身智能与期望的投射。
- 反向图灵测试的概念成为一种合理解释:LLMs如同镜子,揭示了人类而非其自身。
- LLMs本身并非内在智能,但在互动中可能作为放大或反映人类认知模式的工具。
- 将LLMs与感知运动系统结合,是超越纯语言能力、迈向真正自主的关键。
- 提出了七个受大脑启发的改进措施,作为实现LLM系统人工通用自主的路线图。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。