Skip to main content
QUICK REVIEW

[Paper Review] A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia

Queenie Luo, Michael Puett|arXiv (Cornell University)|Mar 28, 2023
Text Readability and Simplification4 citations
TL;DR

This paper investigates language bias in major online platforms—Google, ChatGPT, YouTube, and Wikipedia—demonstrating that search results reflect culturally dominant perspectives tied to the search language, creating a 'perspectival' mirror effect where users see only their own cultural reflections. The study reveals systematic disparities in information coverage across languages, particularly on complex topics like 'Buddhism' and 'colonization,' undermining the promise of diverse, balanced information access.

ABSTRACT

Contrary to Google Search's mission of delivering information from "many angles so you can form your own understanding of the world," we find that Google and its most prominent returned results - Wikipedia and YouTube - simply reflect a narrow set of culturally dominant views tied to the search language for complex topics like "Buddhism," "Liberalism," "colonization," "Iran" and "America." Simply stated, they present, to varying degrees, distinct information across the same search in different languages, a phenomenon we call language bias. This paper presents evidence and analysis of language bias and discusses its larger social implications. We find that our online searches and emerging tools like ChatGPT turn us into the proverbial blind person touching a small portion of an elephant, ignorant of the existence of other cultural perspectives. Language bias sets a strong yet invisible cultural barrier online, where each language group thinks they can see other groups through searches, but in fact, what they see is their own reflection.

Motivation & Objective

  • To investigate how language influences the information retrieved from major online platforms like Google, YouTube, Wikipedia, and ChatGPT.
  • To expose the systemic disparity in content representation across different languages on complex, culturally sensitive topics.
  • To challenge the assumption that search engines provide balanced, multi-perspective views of the world.
  • To highlight how language bias functions as an invisible cultural barrier in digital information access.
  • To contribute to critical discourse on algorithmic fairness and representational equity in AI and search systems.

Proposed method

  • Conducted comparative analysis of search results across multiple languages (e.g., English, Chinese, Arabic, Spanish) for five complex topics: 'Buddhism,' 'Liberalism,' 'colonization,' 'Iran,' and 'America.'
  • Collected and analyzed results from Google Search, YouTube, Wikipedia, and ChatGPT in response to identical queries in different languages.
  • Evaluated content diversity, framing, and representation of cultural perspectives in returned results.
  • Used qualitative and quantitative assessment to compare the depth and scope of information across language groups.
  • Applied a metaphor of the blind men and the elephant to illustrate how each language group perceives only a partial, culturally filtered version of complex topics.
  • Identified patterns of underrepresentation and skewed narratives in non-English results, particularly for non-Western perspectives.

Experimental results

Research questions

  • RQ1How do search results for complex topics vary across different languages on Google, YouTube, Wikipedia, and ChatGPT?
  • RQ2To what extent do these platforms reflect culturally dominant perspectives rather than diverse global viewpoints?
  • RQ3How does language bias affect users' perception of global issues like 'colonization' or 'Liberalism'?
  • RQ4In what ways do AI-generated responses from ChatGPT reflect or amplify language-specific cultural biases?
  • RQ5What are the social implications of language bias in digital information ecosystems?

Key findings

  • Search results for topics like 'Buddhism' and 'colonization' showed significant disparities in content, framing, and cultural representation across languages.
  • Non-English results, particularly in Chinese, Arabic, and Spanish, consistently underrepresented or misrepresented perspectives compared to English results.
  • Wikipedia and YouTube results were heavily influenced by the linguistic and cultural context of the search language, reinforcing dominant narratives.
  • ChatGPT responses in non-English languages exhibited lower factual diversity and higher alignment with culturally dominant viewpoints.
  • The study confirms that users in different language communities access fundamentally different versions of the same topic, creating fragmented, perspective-limited worldviews.
  • Language bias acts as an invisible barrier, where users believe they are accessing global perspectives but are instead seeing reflections of their own cultural context.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.