[Paper Review] Large Language Models Can Infer Personality from Free-Form User Interactions
This study demonstrates that GPT-4-powered chatbots can infer Big Five personality traits from free-form user interactions with moderate accuracy, outperforming static text-based methods. The highest accuracy (mean r = .443) was achieved when chatbots were prompted to elicit personality-relevant information, without compromising user experience, highlighting LLMs' potential for scalable, dynamic psychological profiling in naturalistic settings.
This study investigates the capacity of Large Language Models (LLMs) to infer the Big Five personality traits from free-form user interactions. The results demonstrate that a chatbot powered by GPT-4 can infer personality with moderate accuracy, outperforming previous approaches drawing inferences from static text content. The accuracy of inferences varied across different conversational settings. Performance was highest when the chatbot was prompted to elicit personality-relevant information from users (mean r=.443, range=[.245, .640]), followed by a condition placing greater emphasis on naturalistic interaction (mean r=.218, range=[.066, .373]). Notably, the direct focus on personality assessment did not result in a less positive user experience, with participants reporting the interactions to be equally natural, pleasant, engaging, and humanlike across both conditions. A chatbot mimicking ChatGPT's default behavior of acting as a helpful assistant led to markedly inferior personality inferences and lower user experience ratings but still captured psychologically meaningful information for some of the personality traits (mean r=.117, range=[-.004, .209]). Preliminary analyses suggest that the accuracy of personality inferences varies only marginally across different socio-demographic subgroups. Our results highlight the potential of LLMs for psychological profiling based on conversational interactions. We discuss practical implications and ethical challenges associated with these findings.
Motivation & Objective
- To investigate whether large language models (LLMs) can infer Big Five personality traits from dynamic, free-form user interactions.
- To compare the accuracy of personality inferences across different conversational prompting conditions: personality assessment, naturalistic conversation, and assistant mode.
- To evaluate potential trade-offs between inference accuracy and user experience in different interaction settings.
- To examine demographic biases in LLM-based personality inferences across gender, age, race, education, and socioeconomic status.
- To explore the ethical and practical implications of using LLMs for large-scale, unobtrusive psychological profiling in real-world applications.
Proposed method
- Conducted a 3x2 within-subjects experimental design with 566 U.S. participants on Prolific Academic.
- Used GPT-4 as the LLM-based chatbot under three distinct prompting conditions: personality assessment, naturalistic conversation, and helpful assistant mode.
- Collected free-form conversational interactions where users either engaged naturally or used the chatbot for personal purposes.
- Administered the Big Five Inventory-10 (BFI-10) to participants to obtain ground-truth personality scores for comparison.
- Calculated zero-order correlations (r) between LLM-inferred personality scores and self-reported BFI-10 scores across traits and conditions.
- Conducted user experience assessments using standardized scales for naturalness, pleasantness, engagement, and humanness of interactions.
Experimental results
Research questions
- RQ1Can LLM-based chatbots infer Big Five personality traits from free-form user interactions with higher accuracy than static text-based methods?
- RQ2How does the conversational prompting strategy (e.g., personality assessment vs. naturalistic chat) affect the accuracy of personality inferences?
- RQ3Does prompting the LLM to focus on personality assessment negatively impact user experience compared to more naturalistic or assistant-like interactions?
- RQ4Are there measurable demographic biases in LLM-based personality inferences across gender, age, race, education, and socioeconomic status?
- RQ5To what extent do longer or more naturally occurring interactions improve personality inference accuracy compared to short, task-driven exchanges?
Key findings
- The chatbot prompted to elicit personality-relevant information achieved the highest inference accuracy, with a mean correlation of r = .443 (range: .245 to .640) across the Big Five traits.
- In the naturalistic conversation condition, the mean correlation was r = .218 (range: .066 to .373), indicating moderate but lower accuracy than the targeted assessment condition.
- The default 'helpful assistant' mode yielded the lowest accuracy (mean r = .117, range: -.004 to .209), though still capturing some psychologically meaningful information.
- User experience was consistently high across all conditions, with no significant differences in perceived naturalness, pleasantness, engagement, or humanness.
- Preliminary analyses showed only marginal variation in inference accuracy across demographic subgroups, suggesting limited demographic bias in the model’s performance.
- The study suggests that longer or more naturally occurring interactions could further improve inference accuracy, though current results were based on short interactions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.