[Paper Review] Diversity and Language Technology: How Techno-Linguistic Bias Can Cause Epistemic Injustice
This paper introduces 'techno-linguistic bias'—a systemic, design-level prejudice in language technology that privileges dominant languages and cultures, leading to epistemic injustice by marginalizing non-dominant worldviews. It argues that simply extending AI tools to more languages without addressing underlying design biases perpetuates exclusion and undermines true diversity, advocating for co-creation with marginalized communities to ensure epistemic self-determination.
It is well known that AI-based language technology -- large language models, machine translation systems, multilingual dictionaries, and corpora -- is currently limited to 2 to 3 percent of the world's most widely spoken and/or financially and politically best supported languages. In response, recent research efforts have sought to extend the reach of AI technology to ``underserved languages.'' In this paper, we show that many of these attempts produce flawed solutions that adhere to a hard-wired representational preference for certain languages, which we call techno-linguistic bias. Techno-linguistic bias is distinct from the well-established phenomenon of linguistic bias as it does not concern the languages represented but rather the design of the technologies. As we show through the paper, techno-linguistic bias can result in systems that can only express concepts that are part of the language and culture of dominant powers, unable to correctly represent concepts from other communities. We argue that at the root of this problem lies a systematic tendency of technology developer communities to apply a simplistic understanding of diversity which does not do justice to the more profound differences that languages, and ultimately the communities that speak them, embody. Drawing on the concept of epistemic injustice, we point to the broader sociopolitical consequences of the bias we identify and show how it can lead not only to a disregard for valuable aspects of diversity but also to an under-representation of the needs and diverse worldviews of marginalized language communities.
Motivation & Objective
- To identify and analyze a new form of bias in language technology—techno-linguistic bias—distinct from linguistic bias, rooted in the design and methodology of systems rather than just data or models.
- To demonstrate how this bias results in lexical gaps that exclude culturally specific concepts, such as kinship terms or time systems, from being represented in AI tools.
- To argue that current efforts to expand multilingual AI are insufficient and potentially harmful when they replicate colonial-era hierarchies through one-size-fits-all technological scaling.
- To highlight the sociopolitical consequences of this bias, including the erasure of situated knowledges and the denial of epistemic self-determination for marginalized language communities.
- To call for a fundamental rethinking of language technology development, centered on critical, community-led co-creation that avoids white saviorism and technological paternalism.
Proposed method
- Analyzes existing multilingual language technologies—large language models, machine translation, and corpora—focusing on their design choices and representational limitations.
- Identifies techno-linguistic bias as a by-design preference for languages and cultural frameworks associated with dominant geopolitical powers, especially Anglophone and Western-centric models.
- Uses the concept of epistemic injustice (from Fricker) to frame the exclusion of non-dominant worldviews as a form of systemic knowledge-based oppression.
- Compares language representation across digital platforms (e.g., Wikipedia) to expose disparities between speaker numbers and digital visibility, such as Kiswahili vs. Breton.
- Proposes a multipolar model of language technology development that uses widely spoken trade languages as pivots, but extends it to include critical scrutiny of technological design and power structures.
- Advocates for value-sensitive, community-oriented research methodologies that center the voices and epistemic practices of marginalized language communities in system design.
Experimental results
Research questions
- RQ1How does techno-linguistic bias—embedded in the design of language technologies—differ from linguistic bias and what are its systemic consequences?
- RQ2In what ways do lexical gaps in AI language systems reflect and reproduce the erasure of non-Western worldviews and social practices?
- RQ3Why do current efforts to expand multilingual AI fail to achieve meaningful diversity, and how do they reinforce historical power imbalances?
- RQ4How can language technology development be restructured to avoid epistemic injustice and support epistemic self-determination for marginalized communities?
- RQ5What role do trade languages and multilingual system architectures play in either mitigating or reproducing techno-linguistic bias?
Key findings
- Techno-linguistic bias is a systemic, design-level preference in language technology that privileges certain languages and cultural frameworks, especially those of dominant geopolitical powers, leading to exclusionary outcomes.
- Despite efforts to expand multilingual AI, systems often fail to represent culturally specific concepts—such as kinship relations, time systems, or food terminology—resulting in lexical gaps that reflect epistemic erasure.
- The digital language divide persists not only due to data scarcity but also due to the centralized, top-down development of language technologies that reflect colonial-era power structures.
- The disparity in digital representation is stark: Kiswahili, with 80 million speakers, has far less digital support than Breton, with only 200,000 speakers, due to unequal institutional and technological investment.
- Current AI-based language tools reproduce epistemic injustice by failing to recognize or represent the situated knowledges and worldviews embedded in non-dominant languages.
- The paper concludes that technological expansion without co-creation with language communities risks deepening marginalization and must be restructured around critical, mutual learning frameworks to avoid reproducing colonial power dynamics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.