[Paper Review] Towards Bridging the Digital Language Divide
This paper introduces the LiveLanguage initiative to reduce linguistic bias in multilingual AI by building diversity-aware lexical resources through community-driven, collaborative development. By integrating local languages into a global lexical database (UKC) via participatory methodology, it enables equitable representation and interlingual connectivity, advancing ethical, low-bias language technology design.
It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2-3% most widely spoken languages. Recent research efforts have attempted to expand the coverage of AI technology to `under-resourced languages.' The goal of our paper is to bring attention to a phenomenon that we call linguistic bias: multilingual language processing systems often exhibit a hardwired, yet usually involuntary and hidden representational preference towards certain languages. Linguistic bias is manifested in uneven per-language performance even in the case of similar test conditions. We show that biased technology is often the result of research and development methodologies that do not do justice to the complexity of the languages being represented, and that can even become ethically problematic as they disregard valuable aspects of diversity as well as the needs of the language communities themselves. As our attempt at building diversity-aware language resources, we present a new initiative that aims at reducing linguistic bias through both technological design and methodology, based on an eye-level collaboration with local communities.
Motivation & Objective
- To identify and address linguistic bias in multilingual AI systems, where under-resourced languages suffer from systemic representational disadvantages.
- To challenge the top-down, expert-driven model of language technology development that overlooks local linguistic and cultural complexity.
- To promote linguistic diversity as a core design principle in AI, ensuring that minority and local languages are not marginalized in digital infrastructure.
- To establish a sustainable, community-led model for developing multilingual lexical resources that preserve cultural identity and enable equitable access.
- To demonstrate a scalable, ethically grounded methodology for reducing bias in language technology through collaboration with local institutions and speakers.
Proposed method
- Designing a diversity-aware lexical data model that supports both hub (trade) and satellite (local) languages in a hierarchical multilingual framework.
- Creating a centralized LiveLanguage data catalogue that enables standardized, open access to multilingual lexicons in a common format.
- Implementing a collaborative development pipeline where local institutions lead resource creation, with support from LiveLanguage in tools, training, and infrastructure.
- Integrating local lexicons into the Universal Knowledge Core (UKC), a global lexical database, to enable interlingual mapping and global discovery.
- Providing open-source tools for lexicon management, visualization, and editing, tailored for non-expert users in local communities.
- Establishing a not-for-profit foundation (DataScientia) to ensure long-term sustainability and shared governance of core components.
Experimental results
Research questions
- RQ1How does linguistic bias emerge in multilingual language technology, and what are its root causes in current AI development methodologies?
- RQ2To what extent can community-driven, participatory design reduce linguistic bias in multilingual lexical resources?
- RQ3What technical and methodological frameworks enable equitable integration of under-resourced and local languages into global language infrastructure?
- RQ4How can intellectual property and ownership be managed to empower local institutions while ensuring global interoperability?
- RQ5What role does institutional collaboration play in achieving sustainable, diversity-aware language technology development?
Key findings
- Linguistic bias in AI systems is not accidental but systemic, arising from design choices that privilege dominant languages and overlook cultural and linguistic complexity.
- The LiveLanguage initiative successfully integrated multiple local Alpine languages—such as Mòcheno, Cimbrian, Ladin, and Friulian—into a multilingual lexical framework through community-led development.
- Local institutions retained full intellectual property rights over their contributed lexicons, enabling autonomous dissemination and application of the resources.
- Integration of local lexicons into the UKC database enabled interlingual connectivity, allowing users to access and navigate multilingual data across trade and local languages.
- The project demonstrated that a collaborative, eye-level model with local stakeholders leads to higher-quality, contextually accurate, and ethically sound language resources.
- The establishment of the DataScientia Foundation ensures long-term governance and sustainability of the LiveLanguage ecosystem, with shared decision-making power among diverse stakeholders.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.