Skip to main content
QUICK REVIEW

[论文解读] Diversity and Language Technology: How Techno-Linguistic Bias Can Cause Epistemic Injustice

Paula Helm, Gábor Bella|arXiv (Cornell University)|Jul 25, 2023
Epistemology, Ethics, and Metaphysics被引用 5
一句话总结

本文提出了“技术语言偏见”这一概念——一种系统性、设计层面的语言技术偏见,其优势偏向于主导语言和文化,导致认识论不公,使非主导世界观被边缘化。文章认为,若不解决根本的设计偏见,仅将AI工具扩展至更多语言,只会加剧排斥,破坏真正的多样性,主张通过与边缘化社区共同创造,确保认识论自主权。

ABSTRACT

It is well known that AI-based language technology -- large language models, machine translation systems, multilingual dictionaries, and corpora -- is currently limited to 2 to 3 percent of the world's most widely spoken and/or financially and politically best supported languages. In response, recent research efforts have sought to extend the reach of AI technology to ``underserved languages.'' In this paper, we show that many of these attempts produce flawed solutions that adhere to a hard-wired representational preference for certain languages, which we call techno-linguistic bias. Techno-linguistic bias is distinct from the well-established phenomenon of linguistic bias as it does not concern the languages represented but rather the design of the technologies. As we show through the paper, techno-linguistic bias can result in systems that can only express concepts that are part of the language and culture of dominant powers, unable to correctly represent concepts from other communities. We argue that at the root of this problem lies a systematic tendency of technology developer communities to apply a simplistic understanding of diversity which does not do justice to the more profound differences that languages, and ultimately the communities that speak them, embody. Drawing on the concept of epistemic injustice, we point to the broader sociopolitical consequences of the bias we identify and show how it can lead not only to a disregard for valuable aspects of diversity but also to an under-representation of the needs and diverse worldviews of marginalized language communities.

研究动机与目标

  • 识别并分析语言技术中一种新型偏见——技术语言偏见,其与语言偏见不同,根植于系统的设计与方法论,而非仅数据或模型。
  • 展示此类偏见如何导致词汇空白,使诸如亲属称谓或时间体系等文化特异性概念无法在AI工具中得到体现。
  • 论证当前将多语言AI扩展的努力在未消除殖民时代等级结构的复制下,不仅不足且可能具有危害性,尤其当采用一刀切的技术扩展模式时。
  • 强调此类偏见的社会政治后果,包括地方性知识的抹除,以及对边缘化语言社区认识论自主权的否认。
  • 呼吁对语言技术开发进行根本性反思,以批判性、社区主导的共同创作为中心,避免白人救世主情结与技术家长制。

提出的方法

  • 分析现有的多语言语言技术——大语言模型、机器翻译和语料库——聚焦其设计选择与表征局限。
  • 将技术语言偏见识别为一种有意识的设计偏好,即优先考虑与主导地缘政治权力相关的语言和文化框架,尤其是英语和以西方为中心的模型。
  • 运用弗里克尔(Fricker)提出的“认识论不公”概念,将非主导世界观的排斥视为一种系统性的基于知识的压迫形式。
  • 比较不同数字平台(如维基百科)中的语言表征,揭示使用者数量与数字可见性之间的差异,例如斯瓦希里语与布列塔尼语的对比。
  • 提出一种多极化的语言技术发展模式,以广泛使用的贸易语言为枢纽,但进一步扩展至对技术设计与权力结构的批判性审视。
  • 倡导以价值观为中心、面向社区的研究方法,将边缘化语言社区的声音与认识论实践置于系统设计的中心。

实验结果

研究问题

  • RQ1技术语言偏见——嵌入语言技术设计之中——与语言偏见有何不同?其系统性后果是什么?
  • RQ2AI语言系统中的词汇空白在何种方式上反映并再现了非西方世界观与社会实践的抹除?
  • RQ3为何当前扩展多语言AI的努力无法实现真正意义上的多样性?它们如何强化了历史上的权力失衡?
  • RQ4如何重构语言技术开发,以避免认识论不公,并支持边缘化社区的认识论自主权?
  • RQ5贸易语言与多语言系统架构在缓解或再现技术语言偏见方面发挥何种作用?

主要发现

  • 技术语言偏见是一种系统性、设计层面的偏好,使某些语言和文化框架(尤其是主导地缘政治权力的语言)获得优先地位,导致排斥性结果。
  • 尽管已努力扩展多语言AI,但系统仍常无法体现文化特异性概念(如亲属关系、时间体系或食物术语),导致词汇空白,反映认识论上的抹除。
  • 数字语言鸿沟的持续存在,不仅源于数据稀缺,也源于反映殖民时代权力结构的集中化、自上而下的语言技术开发模式。
  • 数字表征的差异极为显著:斯瓦希里语有8000万使用者,其数字支持远低于仅有20万使用者的布列塔尼语,原因在于制度与技术投资的不平等。
  • 当前基于AI的语言工具通过未能识别或表征非主导语言中嵌入的地方性知识与世界观,再现了认识论不公。
  • 本文结论指出,若无与语言社区的共同创造,单纯的技术扩展将加剧边缘化,必须围绕批判性、相互学习的框架重构,以避免重演殖民权力动态。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。