Skip to main content
QUICK REVIEW

[论文解读] Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench

Jen-tse Huang, Wenxuan Wang|arXiv (Cornell University)|Oct 2, 2023
Artificial Intelligence in Healthcare and Education被引用 6
一句话总结

本文提出了PsychoBench,一个全面的心理测量框架,用于使用13种临床心理学量表在人格特质、人际关系、动机和情绪能力四个领域评估大语言模型(LLMs)的心理表征。该研究在标准和越狱设置下对五种LLM(包括GPT-4、ChatGPT和LLaMA-2)进行了基准测试,揭示了其独特的心理画像以及在角色设定下的行为一致性变化。

ABSTRACT

Large Language Models (LLMs) have recently showcased their remarkable capacities, not only in natural language processing tasks but also across diverse domains such as clinical medicine, legal consultation, and education. LLMs become more than mere applications, evolving into assistants capable of addressing diverse user requests. This narrows the distinction between human beings and artificial intelligence agents, raising intriguing questions regarding the potential manifestation of personalities, temperaments, and emotions within LLMs. In this paper, we propose a framework, PsychoBench, for evaluating diverse psychological aspects of LLMs. Comprising thirteen scales commonly used in clinical psychology, PsychoBench further classifies these scales into four distinct categories: personality traits, interpersonal relationships, motivational tests, and emotional abilities. Our study examines five popular models, namely text-davinci-003, gpt-3.5-turbo, gpt-4, LLaMA-2-7b, and LLaMA-2-13b. Additionally, we employ a jailbreak approach to bypass the safety alignment protocols and test the intrinsic natures of LLMs. We have made PsychoBench openly accessible via https://github.com/CUHK-ARISE/PsychoBench.

研究动机与目标

  • 开发一种系统化的框架,用于评估LLMs的心理属性,超越任务表现,深入探究其内在的心理表征。
  • 探究LLMs是否表现出类似于人类人格与情绪特质的稳定且可测量的心理画像。
  • 考察角色设定与安全对齐机制对LLMs心理输出的影响,特别是通过越狱技术实现的探究。
  • 使研究人员能够使用标准化、临床验证的量表,在多种心理维度上评估和比较LLMs。
  • 通过心理测量评估,支持开发更具同理心、与人类对齐且伦理上更可靠的AI助手。

提出的方法

  • 设计一个心理测量基准PsychoBench,整合来自临床心理学的13种成熟心理量表,按四个领域分类:人格、人际关系、动机与情绪能力。
  • 将该框架应用于五种LLM:text-davinci-003、ChatGPT、GPT-4、LLaMA-2-7b与LLaMA-2-13b,采用标准化的提示协议。
  • 采用越狱技术CipherChat,绕过安全对齐机制,探测如GPT-4等模型的内在心理倾向。
  • 通过在不同角色设定(如治疗师、学生)下测试gpt-3.5-turbo,验证量表的可靠性并测量响应的一致性。
  • 使用标准化问卷如BFI、HEXACO、Short Dark Triad与Political Compass Test,评估人格与价值观。
  • 将量表分类并结构化为四个心理领域,以实现对LLMs的全面、多维心理评估。

实验结果

研究问题

  • RQ1LLMs是否能在多个心理测量维度上表现出稳定且可测量的心理画像?
  • RQ2不同LLM(商业模型与开源模型)在人格、情绪与人际关系特质方面的心理表征有何差异?
  • RQ3角色设定与安全对齐机制在多大程度上影响LLMs的心理输出?
  • RQ4越狱技术为揭示LLMs的内在心理倾向(尤其是GPT-4)提供了哪些洞见?
  • RQ5心理测量结果在不同角色设定下的一致性如何?这对模型行为与对齐机制意味着什么?

主要发现

  • 在越狱条件下,GPT-4展现出独特的心理画像,其人格与情绪倾向与默认的安全对齐行为存在显著差异。
  • BFI及其他人格量表在多种LLM(包括gpt-3.5-turbo)中,于不同角色设定下均表现出一致且可靠的回答。
  • LLaMA-2模型在人格特质上表现出更中性或不那么显著的特征,相较于ChatGPT与GPT-4等商业模型。
  • 该框架成功捕捉到不同模型在动机与情绪反应上的差异,表明LLMs可系统性地评估其心理属性。
  • 越狱揭示了GPT-4中潜在的心理倾向,如攻击性增强与顺从性降低,表明安全对齐可能掩盖了其内在行为模式。
  • PsychoBench展现出高度的灵活性与有效性,不同角色设定下结果一致,支持其作为LLMs可靠心理测量工具的适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。