Skip to main content
QUICK REVIEW

[Paper Review] A Material Lens on Coloniality in NLP

William A. Held, Camille Harris|arXiv (Cornell University)|Nov 14, 2023
Ethics and Social Impacts of AI4 citations
TL;DR

This paper applies Actor-Network Theory (ANT) to expose how colonial power structures are materially embedded in NLP through data, algorithms, and software, revealing increasing geographic and linguistic inequity across NLP research phases. It demonstrates that coloniality is not incidental but systemic, amplified by technological artifacts, and calls for deliberate dismantling of colonial foundations in NLP's data and models to achieve equity.

ABSTRACT

Coloniality, the continuation of colonial harms beyond "official" colonization, has pervasive effects across society and scientific fields. Natural Language Processing (NLP) is no exception to this broad phenomenon. In this work, we argue that coloniality is implicitly embedded in and amplified by NLP data, algorithms, and software. We formalize this analysis using Actor-Network Theory (ANT): an approach to understanding social phenomena through the network of relationships between human stakeholders and technology. We use our Actor-Network to guide a quantitative survey of the geography of different phases of NLP research, providing evidence that inequality along colonial boundaries increases as NLP builds on itself. Based on this, we argue that combating coloniality in NLP requires not only changing current values but also active work to remove the accumulation of colonial ideals in our foundational data and algorithms.

Motivation & Objective

  • To investigate how colonial power dynamics are embedded in the material infrastructure of NLP, including data, annotations, models, and software.
  • To analyze how technological artifacts in NLP reinforce and amplify existing colonial hierarchies across research phases.
  • To quantify the geographic and linguistic disparities in NLP research output from 1950 to 2022, showing increasing inequality along colonial lines.
  • To argue that combating coloniality in NLP requires more than value changes—it demands active deconstruction of colonial ideals embedded in foundational data and algorithms.
  • To use ANT as a formal framework to map and interpret the sociotechnical networks shaping NLP development and deployment.

Proposed method

  • The authors construct a comprehensive Actor-Network of NLP development and deployment, integrating documented processes from prior literature on data production, annotation, modeling, and deployment.
  • They apply Actor-Network Theory (ANT) to treat both human actors and technological artifacts as co-constitutive agents in shaping power dynamics within NLP.
  • A quantitative survey of 10,000+ CL publications up to September 2022 is conducted, using REMIND region classification to map author affiliations and language study origins.
  • The analysis tracks how representation and research focus shift across time, revealing persistent overrepresentation of Western European and settler-colonial institutions.
  • Qualitative case studies are used to reveal limitations of geographic and linguistic metrics, showing that coloniality operates beyond surface-level indicators.
  • The framework is used to interpret how technological systems stabilize colonial power structures even as social and technical conditions evolve.

Experimental results

Research questions

  • RQ1How is coloniality materially embedded in the data, algorithms, and software of NLP systems?
  • RQ2In what ways do technological artifacts in NLP reinforce or reproduce colonial power structures across different research phases?
  • RQ3To what extent does the geographic and linguistic distribution of NLP research reflect colonial-era power imbalances?
  • RQ4How do shifts in geopolitical interest influence the selection of languages studied in NLP, independent of researcher representation?
  • RQ5What role do foundational data and models play in perpetuating colonial ideals, and how can they be actively deconstructed?

Key findings

  • Over 60% of NLP research from 1950 to 2022 was published by institutions in Western Europe and British settler-colonial states, indicating persistent Eurocentric dominance.
  • More than 60% of NLP research focused on Western European languages, reflecting a long-standing linguistic bias rooted in colonial knowledge hierarchies.
  • Surges in research on languages such as Russian and Arabic were driven by geopolitical interests rather than local researcher representation, indicating interest convergence over equity.
  • The Actor-Network analysis reveals that technological artifacts in NLP—especially data and models—act as stabilizers of colonial power, even as social values evolve.
  • Quantitative analysis shows that colonial inequalities increase over time as NLP builds on its own foundational systems, indicating cumulative marginalization.
  • Qualitative case studies demonstrate that geographic and linguistic metrics alone underestimate coloniality, as structural power imbalances persist beneath surface-level data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.