Skip to main content
QUICK REVIEW

[Paper Review] Towards Measuring and Modeling "Culture" in LLMs: A Survey

Muhammad Farid Adilazuarda, Sagnik Mukherjee|arXiv (Cornell University)|Mar 5, 2024
Wikis in Education and Collaboration4 citations
TL;DR

This survey analyzes 39 NLP papers on cultural representation in large language models (LLMs), identifying gaps in defining culture, underexplored cultural proxies like semantic domains and aboutness, and methodological weaknesses in robustness and interpretability. It proposes a framework for categorizing cultural proxies across demographic, semantic, and linguistic-cultural dimensions and calls for interdisciplinary, multilingual, and robust research to achieve holistic cultural inclusion in LLMs.

ABSTRACT

We present a survey of more than 90 recent papers that aim to study cultural representation and inclusion in large language models (LLMs). We observe that none of the studies explicitly define "culture, which is a complex, multifaceted concept; instead, they probe the models on some specially designed datasets which represent certain aspects of "culture". We call these aspects the proxies of culture, and organize them across two dimensions of demographic and semantic proxies. We also categorize the probing methods employed. Our analysis indicates that only certain aspects of ``culture,'' such as values and objectives, have been studied, leaving several other interesting and important facets, especially the multitude of semantic domains (Thompson et al., 2020) and aboutness (Hershcovich et al., 2022), unexplored. Two other crucial gaps are the lack of robustness of probing techniques and situated studies on the impact of cultural mis- and under-representation in LLM-based applications.

Motivation & Objective

  • To analyze how culture is studied in recent NLP research on LLMs, particularly in relation to bias and representation.
  • To identify the lack of a coherent definition of culture across studies and the resulting fragmentation in research.
  • To categorize cultural proxies used in datasets—demographic, semantic, and linguistic-cultural—across 39 surveyed papers.
  • To highlight underexplored cultural dimensions such as aboutness, kinship, time, and physical/mental world concepts.
  • To call for more robust, interpretable, multilingual, and interdisciplinary research to achieve meaningful cultural inclusion in LLMs.

Proposed method

  • Systematically surveyed 39 recent NLP papers on cultural representation, bias, or awareness in LLMs.
  • Categorized cultural proxies into three dimensions: demographic (e.g., region, religion), semantic (e.g., values, norms), and linguistic-cultural (e.g., politeness, speech styles).
  • Analyzed probing methods used in studies, identifying reliance on black-box prompts and lack of white-box interpretability techniques.
  • Evaluated dataset limitations, especially the dominance of English-language data and non-translatability of cultural elements.
  • Proposed a conceptual framework to explicitly link datasets to cultural proxies, enabling clearer research design and reproducibility.
  • Advocated for interdisciplinary engagement with anthropology and HCI to better understand situated cultural contexts in AI.
Towards Measuring and Modeling "Culture" in LLMs: A Survey

Experimental results

Research questions

  • RQ1How is culture defined or operationalized in current NLP research on LLMs?
  • RQ2Which cultural proxies—demographic, semantic, or linguistic-cultural—have been most frequently studied in LLM evaluation?
  • RQ3What are the major gaps in cultural representation, particularly in semantic domains like aboutness, kinship, and time?
  • RQ4To what extent are current probing methods robust and interpretable, and how do they affect the reliability of findings?
  • RQ5How can interdisciplinary and multilingual approaches improve cultural inclusion in LLMs?

Key findings

  • None of the 39 surveyed papers provided a critical or detailed definition of culture, relying instead on vague or high-level descriptions.
  • Only a narrow set of cultural proxies—primarily values and objectives—have been studied, while semantic domains like aboutness, kinship, and function words remain unexplored.
  • Probing methods are largely black-box and sensitive to prompt structure, raising concerns about robustness and generalizability of results.
  • There is a significant lack of multilingual datasets, with most studies relying on English-only data, limiting cross-cultural validity.
  • Interdisciplinary engagement with anthropology and HCI is minimal, despite their relevance for understanding situated and complex cultural dynamics.
  • The survey identifies a critical need for a coherent research framework that explicitly links datasets to cultural proxies and supports robust, interpretable, and inclusive evaluation of LLMs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.