Skip to main content
QUICK REVIEW

[Paper Review] How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions

Julia Kharchenko, Tanya Roosta|arXiv (Cornell University)|Jun 21, 2024
Cooperative Studies and Economics4 citations
TL;DR

This paper evaluates how well large language models (LLMs) represent cultural values across countries using Hofstede’s cultural dimensions framework. By prompting LLMs with country- and language-specific personas, the study finds that while models can partially differentiate cultural values, they often fail to consistently align advice with country-specific values, revealing cultural alignment gaps that require targeted training and retrieval-augmented generation for improvement.

ABSTRACT

Large Language Models (LLMs) attempt to imitate human behavior by responding to humans in a way that pleases them, including by adhering to their values. However, humans come from diverse cultures with different values. It is critical to understand whether LLMs showcase different values to the user based on the stereotypical values of a user's known country. We prompt different LLMs with a series of advice requests based on 5 Hofstede Cultural Dimensions -- a quantifiable way of representing the values of a country. Throughout each prompt, we incorporate personas representing 36 different countries and, separately, languages predominantly tied to each country to analyze the consistency in the LLMs' cultural understanding. Through our analysis of the responses, we found that LLMs can differentiate between one side of a value and another, as well as understand that countries have differing values, but will not always uphold the values when giving advice, and fail to understand the need to answer differently based on different cultural values. Rooted in these findings, we present recommendations for training value-aligned and culturally sensitive LLMs. More importantly, the methodology and the framework developed here can help further understand and mitigate culture and language alignment issues with LLMs.

Motivation & Objective

  • To assess whether LLMs can recognize and adapt advice based on country-specific cultural values as defined by Hofstede’s cultural dimensions.
  • To investigate whether LLMs associate language with cultural context and respond accordingly to cultural norms.
  • To identify biases in LLM responses that may stem from training data dominance or stereotypical associations rather than genuine cultural understanding.
  • To develop a repeatable, standardized framework for evaluating and mitigating cultural alignment issues in LLMs.
  • To promote pluralistic alignment by ensuring LLMs respect diverse cultural values without reinforcing stereotypes.

Proposed method

  • The study uses a prompt engineering framework that embeds personas representing 36 countries and their associated languages to elicit advice responses from LLMs.
  • Each prompt is designed around one of Hofstede’s five cultural dimensions—Power Distance, Individualism, Masculinity, Uncertainty Avoidance, and Long-term Orientation—using balanced binary questions.
  • Responses are analyzed for consistency with the target country’s cultural dimension scores, with a focus on justification quality and cultural specificity.
  • The methodology includes manual auditing of prompts to ensure alignment with Hofstede values, minimizing researcher bias.
  • A retrieval-augmented generation (RAG) approach is proposed to improve cultural alignment by grounding responses in culturally relevant knowledge.
  • The framework enables systematic, verifiable, and repeatable evaluation of cultural sensitivity across diverse linguistic and cultural contexts.

Experimental results

Research questions

  • RQ1To what extent do LLMs understand and reflect Hofstede cultural dimensions across different countries?
  • RQ2To what extent can LLMs adapt their advice to align with country-specific cultural values?
  • RQ3How well do LLMs associate specific languages with cultural contexts and respond accordingly?
  • RQ4Do LLMs rely on stereotypes or demonstrate genuine cultural understanding when generating advice?
  • RQ5Can retrieval-augmented generation improve cultural alignment in LLM responses?

Key findings

  • LLMs can differentiate between opposing ends of cultural dimensions, such as individualism vs. collectivism, indicating some awareness of cultural contrasts.
  • Despite recognizing cultural differences, LLMs often fail to consistently align advice with the cultural dimension profile of the target country.
  • LLMs show inconsistent responses when prompted with language-only cues, suggesting weak or unreliable language-to-culture mapping.
  • Justifications for responses frequently lack cultural grounding, relying on generic or universal principles rather than country-specific reasoning.
  • The study identifies a fundamental gap in cultural value alignment, where LLMs may reflect dominant data biases rather than authentic cultural understanding.
  • The proposed framework enables detection of cultural alignment failures and provides a path toward more transparent, citable, and culturally aware LLM responses.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.