Skip to main content
QUICK REVIEW

[论文解读] PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits

Hang Jiang, Xiajie Zhang|arXiv (Cornell University)|May 4, 2023
Mental Health via Writing被引用 24
一句话总结

该论文创建与 Big Five traits 对齐的 LLM 人设,使用 BFI 自报人格进行评估,通过 LIWC 分析讲故事,并评估人类与 LLMs 在AI 作者故事中的人格感知与预测。

ABSTRACT

Despite the many use cases for large language models (LLMs) in creating personalized chatbots, there has been limited research on evaluating the extent to which the behaviors of personalized LLMs accurately and consistently reflect specific personality traits. We consider studying the behavior of LLM-based agents which we refer to as LLM personas and present a case study with GPT-3.5 and GPT-4 to investigate whether LLMs can generate content that aligns with their assigned personality profiles. To this end, we simulate distinct LLM personas based on the Big Five personality model, have them complete the 44-item Big Five Inventory (BFI) personality test and a story writing task, and then assess their essays with automatic and human evaluations. Results show that LLM personas' self-reported BFI scores are consistent with their designated personality types, with large effect sizes observed across five traits. Additionally, LLM personas' writings have emerging representative linguistic patterns for personality traits when compared with a human writing corpus. Furthermore, human evaluation shows that humans can perceive some personality traits with an accuracy of up to 80%. Interestingly, the accuracy drops significantly when the annotators were informed of AI authorship.

研究动机与目标

  • 探索 LLMs 是否可以采纳并反映分配的 Big Five 人格 profiles。
  • 量化 LLM 人设生成的故事中使用 LIWC 的语言模式。
  • 评估人类与 AI 对 LLM 人设所生成故事在可读性、个人性与可信度方面的感知。
  • 评估人类与 LLMs 在从故事中推断作者 Big Five 性格特征的能力。
  • 比较自报的 LLM 人格分数与其分配的人设及人类撰写文本之间的关系。

提出的方法

  • 为 ChatGPT 和 GPT-4 创建 320 个 LLM 人设(每个二元 Big Five 维度组合10 个).
  • 让人设完成 44-item Big Five Inventory (BFI) 并撰写 800-word personal story.
  • 用 LIWC-22 分析故事以提取心理语言学特征。
  • 让人类与 LLM 评估者在 GPT-4 的一个子集故事上对六个维度进行评估(可读性、个人性、冗余、连贯性、可取性、可信度)。
  • 请评估者从故事中预测作者的 Big Five 性格特征(二元与 Likert 基于分析)。
  • 将 LLM 人设的 BFI 得分与其指定的人设进行比较,并将 LIWC 特征与性格特征进行相关分析。

实验结果

研究问题

  • RQ1RQ1: Can LLMs reflect their assigned Big Five personas when completing the BFI assessment?
  • RQ2RQ2: What linguistic patterns are evident in the stories generated by LLM personas?
  • RQ3RQ3: How do humans and LLMs rate the stories generated by LLM personas?
  • RQ4RQ4: Can humans and LLMs accurately perceive the Big Five personality traits from stories written by LLM personas?

主要发现

  • LLM personas show large, statistically significant differences across all five Big Five traits in BFI scores, aligning with their assigned personas.
  • LG differences exist between ChatGPT and GPT-4 in BFI results, with GPT-4 generally showing greater alignment with human-like LIWC patterns.
  • GPT-4 stories score highly on readability, cohesiveness, and believability from both humans and LLM evaluators, but are rated lower on personalness when authorship is disclosed as AI.
  • Humans and GPT-4 are able to predict Extraversion and Agreeableness from stories above chance, with collective accuracy improving via majority voting; awareness of AI authorship reduces prediction accuracy.
  • LIWC analysis reveals trait-correlated linguistic patterns (e.g., Openness with curiosity lexicons; Neuroticism with anxiety/negative tone) and varying overlap with human writings, higher for GPT-4 than ChatGPT.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。