[论文解读] A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia
本文研究了谷歌、ChatGPT、YouTube 和维基百科等主要在线平台的语言偏见,表明搜索结果反映了与搜索语言相关的文化主导视角,形成一种‘视角化’的镜像效应,使用户仅看到自身文化的映射。研究揭示了在复杂话题(如‘佛教’和‘殖民主义’)上,不同语言间信息覆盖存在系统性差异,削弱了获取多样化、平衡信息的承诺。
Contrary to Google Search's mission of delivering information from "many angles so you can form your own understanding of the world," we find that Google and its most prominent returned results - Wikipedia and YouTube - simply reflect a narrow set of culturally dominant views tied to the search language for complex topics like "Buddhism," "Liberalism," "colonization," "Iran" and "America." Simply stated, they present, to varying degrees, distinct information across the same search in different languages, a phenomenon we call language bias. This paper presents evidence and analysis of language bias and discusses its larger social implications. We find that our online searches and emerging tools like ChatGPT turn us into the proverbial blind person touching a small portion of an elephant, ignorant of the existence of other cultural perspectives. Language bias sets a strong yet invisible cultural barrier online, where each language group thinks they can see other groups through searches, but in fact, what they see is their own reflection.
研究动机与目标
- 调查语言如何影响从谷歌、YouTube、维基百科和 ChatGPT 等主要在线平台检索到的信息。
- 揭示在复杂、具有文化敏感性的主题上,不同语言间内容呈现的系统性差异。
- 挑战搜索引擎能提供世界多重视角、平衡观点的假设。
- 突出语言偏见如何作为数字信息获取中的无形文化障碍。
- 为人工智能和搜索系统中的算法公平性与代表性公平性问题,贡献批判性讨论。
提出的方法
- 对五个复杂主题(‘佛教’、‘自由主义’、‘殖民主义’、‘伊朗’和‘美国’)在多种语言(如英语、中文、阿拉伯语、西班牙语)中的搜索结果进行对比分析。
- 针对不同语言中的相同查询,收集并分析谷歌搜索、YouTube、维基百科和 ChatGPT 的结果。
- 评估返回结果中的内容多样性、框架方式以及文化视角的呈现。
- 采用定性与定量评估方法,比较不同语言群体间信息的深度与广度。
- 运用‘盲人摸象’的隐喻,说明每个语言群体仅能感知复杂主题的部分、受文化过滤的版本。
- 识别出非英语结果中,特别是非西方视角的代表性不足与叙事偏移的模式。
实验结果
研究问题
- RQ1谷歌、YouTube、维基百科和 ChatGPT 上,复杂主题的搜索结果在不同语言间如何变化?
- RQ2这些平台在多大程度上反映的是文化主导视角,而非多样化的全球观点?
- RQ3语言偏见在多大程度上影响用户对‘殖民主义’或‘自由主义’等全球议题的认知?
- RQ4ChatGPT 生成的非英语回答在多大程度上反映或放大了特定语言的文化偏见?
- RQ5语言偏见在数字信息生态系统中具有哪些社会影响?
主要发现
- ‘佛教’和‘殖民主义’等主题的搜索结果在内容、框架和文化呈现方面,不同语言间存在显著差异。
- 非英语结果,尤其是中文、阿拉伯语和西班牙语结果,与英语结果相比,持续存在代表性不足或错误呈现的问题。
- 维基百科和 YouTube 的结果受搜索语言的语言和文化背景强烈影响,强化了主导叙事。
- ChatGPT 在非英语语言中的回答事实多样性较低,且更倾向于与文化主导观点保持一致。
- 本研究证实,不同语言社群的用户所接触的同一主题的根本版本截然不同,导致碎片化、视角受限的世界观。
- 语言偏见构成了一道无形障碍,用户误以为自己在获取全球视角,实则看到的只是自身文化语境的映射。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。