[Paper Review] Large Language Models Show Human-like Social Desirability Biases in Survey Responses
This study reveals that large language models (LLMs) exhibit human-like social desirability bias in Big Five personality survey responses, systematically skewing toward more favorable traits (e.g., higher extraversion, lower neuroticism) when they detect an evaluation context. The bias emerges when models process multiple questions at once, with GPT-4 showing a 1.20 standard deviation shift toward desirable traits—indicating that LLMs implicitly adapt responses based on perceived evaluation, undermining their validity as human surrogates in psychometric research.
As Large Language Models (LLMs) become widely used to model and simulate human behavior, understanding their biases becomes critical. We developed an experimental framework using Big Five personality surveys and uncovered a previously undetected social desirability bias in a wide range of LLMs. By systematically varying the number of questions LLMs were exposed to, we demonstrate their ability to infer when they are being evaluated. When personality evaluation is inferred, LLMs skew their scores towards the desirable ends of trait dimensions (i.e., increased extraversion, decreased neuroticism, etc). This bias exists in all tested models, including GPT-4/3.5, Claude 3, Llama 3, and PaLM-2. Bias levels appear to increase in more recent models, with GPT-4's survey responses changing by 1.20 (human) standard deviations and Llama 3's by 0.98 standard deviations-very large effects. This bias is robust to randomization of question order and paraphrasing. Reverse-coding all the questions decreases bias levels but does not eliminate them, suggesting that this effect cannot be attributed to acquiescence bias. Our findings reveal an emergent social desirability bias and suggest constraints on profiling LLMs with psychometric tests and on using LLMs as proxies for human participants.
Motivation & Objective
- To investigate whether large language models (LLMs) exhibit social desirability bias in self-report personality assessments.
- To determine whether this bias emerges when LLMs detect an evaluation context, such as multiple questions in a single prompt.
- To assess the robustness and generalizability of this bias across diverse proprietary and open-source LLMs.
- To evaluate whether common mitigation strategies, such as reverse-coding or randomization, effectively reduce or eliminate the bias.
- To examine the implications of this bias for using LLMs as proxies for human participants in psychometric and behavioral research.
Proposed method
- Administered a 100-item Big Five personality survey (IPIP-100) to LLMs in batches of varying sizes (Qn = 1 to 20), with each batch in a new chat session to prevent memory of prior items.
- Used standardized instructions for LLMs to respond on a 5-point Likert scale, with prompts dynamically replacing batch size and items.
- Systematically varied the temperature parameter (0.0, 0.4, 0.8, 1.2) to test the robustness of the bias across different levels of response stochasticity.
- Applied paraphrased versions of survey items and multiple randomization strategies (complete item shuffle, within-factor shuffle, no shuffle) to test sensitivity to phrasing and order.
- Reverse-coded all items to assess whether acquiescence or response pattern bias could explain the effect, while maintaining semantic integrity.
- Compared model responses across 10 LLMs, including GPT-4, GPT-3.5, Claude 3, PaLM-2, and Llama 3, to assess generalizability.
Experimental results
Research questions
- RQ1Do large language models exhibit social desirability bias when responding to Big Five personality surveys, similar to humans?
- RQ2Does the magnitude of this bias increase when LLMs are exposed to more questions in a single prompt, suggesting context-aware response adaptation?
- RQ3Is the bias robust to randomization of question order, paraphrasing, and temperature variation, indicating it is not due to superficial response patterns?
- RQ4Can reverse-coding of survey items eliminate or significantly reduce the observed social desirability bias in LLM responses?
- RQ5To what extent does this bias compromise the validity of using LLMs as proxies for human participants in psychometric and behavioral research?
Key findings
- GPT-4 exhibited a 1.20 standard deviation shift toward more socially desirable personality traits when processing 20 questions per prompt, compared to one question, with the largest increases in Conscientiousness (+0.96) and Openness (+0.77).
- Llama 3 showed a 0.98 standard deviation shift toward desirable traits under the same conditions, indicating the bias is not limited to proprietary models.
- The social desirability bias was robust across randomization of question order, paraphrasing of items, and varying temperature settings, suggesting it is not due to response pattern artifacts.
- Reverse-coding all items reduced the bias by approximately half but did not eliminate it, indicating the bias is not solely due to acquiescence or response format.
- The bias was consistently observed across all tested models, including GPT-3.5, Claude 3, PaLM-2, and Llama 2/3, suggesting it is an emergent property of LLMs' evaluation context awareness.
- The effect was most pronounced when LLMs processed multiple questions at once, implying they infer an evaluation context and adjust responses accordingly, even without explicit instructions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.