[Paper Review] PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
The paper creates LLM personas aligned with the Big Five traits, assesses their self-reported personalities via BFI, analyzes their storytelling with LIWC, and evaluates how humans and LLMs perceive and predict personality from AI-authored stories.
Despite the many use cases for large language models (LLMs) in creating personalized chatbots, there has been limited research on evaluating the extent to which the behaviors of personalized LLMs accurately and consistently reflect specific personality traits. We consider studying the behavior of LLM-based agents which we refer to as LLM personas and present a case study with GPT-3.5 and GPT-4 to investigate whether LLMs can generate content that aligns with their assigned personality profiles. To this end, we simulate distinct LLM personas based on the Big Five personality model, have them complete the 44-item Big Five Inventory (BFI) personality test and a story writing task, and then assess their essays with automatic and human evaluations. Results show that LLM personas' self-reported BFI scores are consistent with their designated personality types, with large effect sizes observed across five traits. Additionally, LLM personas' writings have emerging representative linguistic patterns for personality traits when compared with a human writing corpus. Furthermore, human evaluation shows that humans can perceive some personality traits with an accuracy of up to 80%. Interestingly, the accuracy drops significantly when the annotators were informed of AI authorship.
Motivation & Objective
- Explore whether LLMs can adopt and reflect assigned Big Five personality profiles.
- Quantify linguistic patterns in stories generated by LLM personas using LIWC.
- Assess human and AI perceptions of AI-generated stories on readability, personalness, and believability.
- Evaluate how well humans and LLMs can infer the writer's Big Five traits from stories.
- Compare self-reported LLM personality scores with their assigned personas and human-written texts.
Proposed method
- Create 320 LLM personas (10 per each binary Big Five trait combination) for ChatGPT and GPT-4.
- Have personas complete the 44-item Big Five Inventory (BFI) and write an 800-word personal story.
- Analyze stories with LIWC-22 to extract psycholinguistic features.
- Have human and LLM raters evaluate a subset of GPT-4 stories on six dimensions (readability, personalness, redundancy, cohesiveness, likeability, believability).
- Ask raters to predict the writer’s Big Five traits from stories (binary and Likert-based analyses).
- Compare LLM persona BFI scores to their designated profiles and correlate LIWC features with personality traits.
Experimental results
Research questions
- RQ1RQ1: Can LLMs reflect their assigned Big Five personas when completing the BFI assessment?
- RQ2RQ2: What linguistic patterns are evident in the stories generated by LLM personas?
- RQ3RQ3: How do humans and LLMs rate the stories generated by LLM personas?
- RQ4RQ4: Can humans and LLMs accurately perceive the Big Five personality traits from stories written by LLM personas?
Key findings
- LLM personas show large, statistically significant differences across all five Big Five traits in BFI scores, aligning with their assigned personas.
- LG differences exist between ChatGPT and GPT-4 in BFI results, with GPT-4 generally showing greater alignment with human-like LIWC patterns.
- GPT-4 stories score highly on readability, cohesiveness, and believability from both humans and LLM evaluators, but are rated lower on personalness when authorship is disclosed as AI.
- Humans and GPT-4 are able to predict Extraversion and Agreeableness from stories above chance, with collective accuracy improving via majority voting; awareness of AI authorship reduces prediction accuracy.
- LIWC analysis reveals trait-correlated linguistic patterns (e.g., Openness with curiosity lexicons; Neuroticism with anxiety/negative tone) and varying overlap with human writings, higher for GPT-4 than ChatGPT.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.