[Paper Review] Can ChatGPT Read Who You Are?
This study investigates whether ChatGPT can infer personality traits from short Czech texts using the Big Five Inventory (BFI) as ground truth. It demonstrates that ChatGPT performs competitively against human raters, though it exhibits a consistent positivity bias across all traits, with performance significantly influenced by prompt design.
The interplay between artificial intelligence (AI) and psychology, particularly in personality assessment, represents an important emerging area of research. Accurate personality trait estimation is crucial not only for enhancing personalization in human-computer interaction but also for a wide variety of applications ranging from mental health to education. This paper analyzes the capability of a generic chatbot, ChatGPT, to effectively infer personality traits from short texts. We report the results of a comprehensive user study featuring texts written in Czech by a representative population sample of 155 participants. Their self-assessments based on the Big Five Inventory (BFI) questionnaire serve as the ground truth. We compare the personality trait estimations made by ChatGPT against those by human raters and report ChatGPT's competitive performance in inferring personality traits from text. We also uncover a 'positivity bias' in ChatGPT's assessments across all personality dimensions and explore the impact of prompt composition on accuracy. This work contributes to the understanding of AI capabilities in psychological assessment, highlighting both the potential and limitations of using large language models for personality inference. Our research underscores the importance of responsible AI development, considering ethical implications such as privacy, consent, autonomy, and bias in AI applications.
Motivation & Objective
- To evaluate the capability of ChatGPT to infer personality traits from short written texts in a non-English language (Czech).
- To compare ChatGPT’s personality trait estimations against those of human raters using self-reported BFI assessments as ground truth.
- To investigate the presence and impact of a positivity bias in LLM-based personality assessments.
- To analyze how different prompt compositions affect the accuracy of personality inference by ChatGPT.
- To highlight ethical implications such as privacy, consent, autonomy, and bias in AI-driven user modeling.
Proposed method
- Conducted a user study with 155 Czech-speaking participants who wrote four types of short texts and completed the Big Five Inventory (BFI) for personality assessment.
- Collected self-assessments from participants as the ground truth reference for personality traits across five dimensions: extraversion, agreeableness, conscientiousness, neuroticism, and openness.
- Employed multiple prompt variations to assess ChatGPT’s performance, including zero-shot, few-shot, and chain-of-thought prompting strategies.
- Used statistical descriptives (mean, median, mode, standard deviation, coefficient of variation) to evaluate and compare the distribution of ChatGPT’s predictions against human and self-assessments.
- Analyzed the consistency and accuracy of ChatGPT’s predictions across three levels (low, neutral, high) of each personality dimension.
- Compared ChatGPT’s outputs with those of human raters (H_A and H_B) and other LLM variants (GPT-TL, GPT-DTL, etc.) to assess relative performance and bias patterns.

Experimental results
Research questions
- RQ1Can ChatGPT accurately infer personality traits from short texts written in Czech, as measured against self-reported BFI assessments?
- RQ2How does ChatGPT’s performance in personality inference compare to that of human raters?
- RQ3Does ChatGPT exhibit a systematic positivity bias in its personality trait estimations across all five dimensions?
- RQ4To what extent does prompt composition influence the accuracy of personality inference by ChatGPT?
- RQ5What are the ethical implications of using LLMs like ChatGPT for automatic personality assessment in real-world applications?
Key findings
- ChatGPT demonstrated competitive performance in inferring personality traits from short Czech texts, with mean scores across dimensions closely aligning with self-assessments in many cases.
- A consistent positivity bias was observed in ChatGPT’s assessments: for all five personality traits, predicted scores were systematically higher than self-assessments, especially in the low and neutral ranges.
- The mean prediction for neuroticism was 1.805 in the high self-assessment group, indicating a strong tendency to underestimate high neuroticism, while mean scores for openness in the high group were 2.041, suggesting underestimation of high openness.
- Prompt design significantly influenced performance: few-shot and chain-of-thought prompting variants (e.g., GPT-DLT) showed improved accuracy over zero-shot baselines, with GPT-DLT achieving a mean of 2.789 for openness in the high self-assessment group.
- Human raters (H_A and H_B) showed higher consistency with self-assessments than ChatGPT in most dimensions, particularly in neuroticism and openness, where ChatGPT’s mean predictions were consistently lower than self-reports.
- The coefficient of variation for ChatGPT’s predictions was generally lower than for human raters, indicating more consistent but less variable responses, which may reflect a lack of nuanced sensitivity to individual differences.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.