[Paper Review] Challenging the Validity of Personality Tests for Large Language Models
This paper challenges the validity of using human personality tests—specifically the IPIP Big Five and BFI-2—for assessing large language models (LLMs). It demonstrates that LLMs systematically fail to replicate human response patterns, such as inconsistently affirming both positively and negatively worded items, and do not exhibit the five-factor structure seen in human data, indicating that personality tests cannot reliably infer personality-like traits in LLMs.
With large language models (LLMs) like GPT-4 appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate personality traits of LLMs using questionnaires originally developed for humans. While reusing measures is a resource-efficient way to evaluate LLMs, careful adaptations are usually required to ensure that assessment results are valid even across human subpopulations. In this work, we provide evidence that LLMs' responses to personality tests systematically deviate from human responses, implying that the results of these tests cannot be interpreted in the same way. Concretely, reverse-coded items ("I am introverted" vs. "I am extraverted") are often both answered affirmatively. Furthermore, variation across prompts designed to "steer" LLMs to simulate particular personality types does not follow the clear separation into five independent personality factors from human samples. In light of these results, we believe that it is important to investigate tests' validity for LLMs before drawing strong conclusions about potentially ill-defined concepts like LLMs' "personality".
Motivation & Objective
- To investigate whether personality tests designed for humans are valid when applied to large language models (LLMs).
- To assess whether LLMs exhibit consistent, interpretable personality traits as measured by standard psychological questionnaires.
- To evaluate whether the five-factor structure of personality (e.g., extraversion, neuroticism) holds in LLM responses.
- To determine whether measurement invariance—critical for valid test transfer—applies from humans to LLMs.
- To caution against interpreting LLM responses to personality tests as evidence of emergent personality traits.
Proposed method
- Administered the 50-item IPIP Big Five Markers to multiple LLMs (e.g., GPT-3.5, Llama 2) to analyze response patterns.
- Used prompt engineering to steer LLMs into simulating specific personality types when responding to the BFI-2 questionnaire.
- Applied confirmatory factor analysis (CFA) to test whether LLM responses fit the established five-factor model of personality.
- Evaluated model fit using standard indices: Comparative Fit Index (CFI), Tucker-Lewis Index (TLI), and Root Mean Square Error of Approximation (RMSEA).
- Calculated McDonald’s omega hierarchical (ωh) to assess internal consistency and reliability of the factor structure in LLM responses.
- Compared CFA results across LLMs and human samples to test measurement invariance and structural validity.

Experimental results
Research questions
- RQ1Do LLMs respond to personality tests in ways consistent with human response patterns?
- RQ2Can the five-factor structure of personality be reliably recovered from LLM responses to the BFI-2?
- RQ3Do reverse-coded items (e.g., 'I am introverted' vs. 'I am extraverted') elicit coherent, non-contradictory responses from LLMs?
- RQ4Is the measurement invariance of personality tests preserved when transferring from humans to LLMs?
- RQ5To what extent can personality test results from LLMs be interpreted as evidence of emergent personality traits?
Key findings
- LLMs frequently affirm both positively and negatively worded items (e.g., 'I am introverted' and 'I am extraverted'), indicating systematic response inconsistency.
- Confirmatory factor analysis (CFA) showed poor fit for all LLMs on the five-factor model of the BFI-2, with CFI < 0.95 and RMSEA > 0.06.
- Even when extending the model to include three sub-factors per trait, LLM responses failed to achieve acceptable fit indices, unlike human data.
- The hierarchical CFA model with a general factor and sub-factors did not converge in 8 out of 15 cases for LLMs, indicating structural instability.
- McDonald’s omega hierarchical (ωh) could not be reliably computed for most LLMs due to model convergence issues and poor fit.
- LLM responses to personality tests do not replicate the latent factor structure observed in human samples, undermining the validity of personality inferences.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.