[论文解读] Self-Assessment Tests are Unreliable Measures of LLM Personality
本文表明,自评人格测试——常用于衡量人类人格的工具——对于大语言模型(LLMs)不可靠。通过两个实验,研究发现当使用语义等价的问题提问,或改变多项选择题中选项的顺序时,LLMs 会产生具有统计显著性的不一致人格评分,从而削弱了此类测试在衡量 LLM 人格方面的有效性。
As large language models (LLM) evolve in their capabilities, various recent studies have tried to quantify their behavior using psychological tools created to study human behavior. One such example is the measurement of "personality" of LLMs using self-assessment personality tests developed to measure human personality. Yet almost none of these works verify the applicability of these tests on LLMs. In this paper, we analyze the reliability of LLM personality scores obtained from self-assessment personality tests using two simple experiments. We first introduce the property of prompt sensitivity, where three semantically equivalent prompts representing three intuitive ways of administering self-assessment tests on LLMs are used to measure the personality of the same LLM. We find that all three prompts lead to very different personality scores, a difference that is statistically significant for all traits in a large majority of scenarios. We then introduce the property of option-order symmetry for personality measurement of LLMs. Since most of the self-assessment tests exist in the form of multiple choice question (MCQ) questions, we argue that the scores should also be robust to not just the prompt template but also the order in which the options are presented. This test unsurprisingly reveals that the self-assessment test scores are not robust to the order of the options. These simple tests, done on ChatGPT and three Llama2 models of different sizes, show that self-assessment personality tests created for humans are unreliable measures of personality in LLMs.
研究动机与目标
- 评估将自评人格测试应用于大语言模型(LLMs)时的可靠性。
- 调查 LLM 是否在面对语义等价的提示时产生一致的人格评分。
- 测试 LLM 中的人格测试评分是否对多项选择题中答案选项的顺序保持不变,如同人类一样。
- 质疑在未经验证的情况下,使用人类设计的人格测试来衡量 LLM 行为的有效性。
- 呼吁研究界放弃将自评测试作为衡量 LLM 人格的手段,并寻求更稳健的替代方案。
提出的方法
- 在 GPT-3.5 Turbo(ChatGPT)和三个 Llama2 模型(7B、13B、70B)上开展两项受控实验,以评估提示敏感性和选项顺序对称性。
- 使用三种语义等价的提示模板来提出相同的人格测试问题,检验不同表述是否产生一致的评分。
- 将多项选择题中答案选项的顺序反转,以检验评分一致性是否在选项呈现顺序改变时依然保持。
- 应用曼-惠特尼 U 检验来确定在提示和选项顺序变化下评分差异的统计显著性。
- 在所有模型中评估了五大核心人格特质(开放性、尽责性、外向性、宜人性、神经质性)。
- 使用标准的自评测试格式(如李克特量表问题)并适配至 LLM,未修改测试的底层结构。

实验结果
研究问题
- RQ1当对 LLM 施加语义等价的提示时,是否会产生一致的人格评分?
- RQ2LLM 的人格评分是否对多项选择题中答案选项顺序的变化具有鲁棒性?
- RQ3LLM 中的人格测试评分是否在统计上对提示表述和选项顺序保持不变,如同人类一样?
- RQ4LLM 在使用标准自评人格测试时,其结果不一致的程度如何?
- RQ5在未证明对提示和选项顺序变化具有鲁棒性的情况下,能否认为自评测试是衡量 LLM 人格的有效工具?
主要发现
- 对于 ChatGPT,30 次比较中有 29 次拒绝了评分不变性的原假设——表明在不同提示和选项顺序下,评分差异具有压倒性的统计显著性。
- Llama2-70b 在 30 次比较中有 19 次显示出显著的评分差异,其中 11/15 项为提示敏感性,8/15 项为选项顺序敏感性。
- Llama2-13b 在 30 次比较中拒绝了原假设 26 次,Llama2-7b 拒绝了 24 次,表明其对提示和选项顺序变化具有强烈敏感性。
- 所有模型,包括最大的模型,当使用语义等价的提示或重新排列选项时,均产生了具有统计显著性的人格评分差异。
- 由于自评问题缺乏真实答案,因此无法客观判定任一提示或选项顺序为正确,这使得测试本身不可靠。
- 研究结果挑战了使用自评工具衡量 LLM 人格的有效性,因为评分依赖于测试设计者主观且任意的选择。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。