Skip to main content
QUICK REVIEW

[Paper Review] Large Language Models Can Infer Psychological Dispositions of Social Media Users

Heinrich Peters, Sandra Matz|arXiv (Cornell University)|Sep 13, 2023
Artificial Intelligence in Healthcare and Education13 citations
TL;DR

The study shows that GPT-3.5 and GPT-4 can infer Big Five personality traits from Facebook status updates in a zero-shot setting, achieving an average correlation around .29 with self-reports and exhibiting gender and age biases.

ABSTRACT

Large Language Models (LLMs) demonstrate increasingly human-like abilities across a wide variety of tasks. In this paper, we investigate whether LLMs like ChatGPT can accurately infer the psychological dispositions of social media users and whether their ability to do so varies across socio-demographic groups. Specifically, we test whether GPT-3.5 and GPT-4 can derive the Big Five personality traits from users' Facebook status updates in a zero-shot learning scenario. Our results show an average correlation of r = .29 (range = [.22, .33]) between LLM-inferred and self-reported trait scores - a level of accuracy that is similar to that of supervised machine learning models specifically trained to infer personality. Our findings also highlight heterogeneity in the accuracy of personality inferences across different age groups and gender categories: predictions were found to be more accurate for women and younger individuals on several traits, suggesting a potential bias stemming from the underlying training data or differences in online self-expression. The ability of LLMs to infer psychological dispositions from user-generated text has the potential to democratize access to cheap and scalable psychometric assessments for both researchers and practitioners. On the one hand, this democratization might facilitate large-scale research of high ecological validity and spark innovation in personalized services. On the other hand, it also raises ethical concerns regarding user privacy and self-determination, highlighting the need for stringent ethical frameworks and regulation.

Motivation & Objective

  • Assess whether LLMs can infer the Big Five personality traits from social media text without explicit training.
  • Evaluate zero-shot inference performance of GPT-3.5 and GPT-4 using Facebook status updates.
  • Examine potential demographic biases (gender and age) in LLM-based inferences.

Proposed method

  • Use 1000 MyPersonality participants with IPIP self-reports and at least 200 Facebook status updates.
  • Concatenate the last 200 status updates per user and prompt GPT-3.5 and GPT-4 to rate Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism on a 1–5 scale.
  • Process updates in 20-message chunks and average three rating rounds to obtain aggregate trait scores.
  • Compare LLM-inferred scores with self-reported IPIP scores using Pearson correlations.
  • Assess accuracy differences across gender and age groups via residual analyses.

Experimental results

Research questions

  • RQ1Can GPT-3.5 and GPT-4 infer the Big Five personality traits from social media text in a zero-shot setting?
  • RQ2How do the inferred traits correlate with self-reported trait scores, and how does this vary by trait and model version?
  • RQ3Do gender or age affect the accuracy or bias of LLM-based personality inferences?
  • RQ4How does the amount of input text influence inference accuracy?

Key findings

  • GPT-3.5 achieved an average correlation of r = .27; GPT-4 achieved r = .31 across all traits.
  • Trait-level correlations were highest for Openness (.28 / .33), Extraversion (.29 / .32), and Agreeableness (.30 / .32) for GPT-3.5 / GPT-4 respectively.
  • Conscientiousness showed lower correlations (.22 / .26) and Neuroticism (.26 / .29) for GPT-3.5 / GPT-4 respectively.
  • Overall, GPT-4 provided more accurate inferences than GPT-3.5, though none of the trait-wise differences were statistically significant after correction.
  • Gender analyses suggest women show higher inferred scores on several traits and that male inferences have larger residuals, indicating lower accuracy for men in multiple traits.
  • Age analyses indicate older users show higher self-reported Conscientiousness and Neuroticism-related differences, with some traits showing reduced accuracy for older users depending on the model.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.