Skip to main content
QUICK REVIEW

[Paper Review] The Limits of Global Inclusion in AI Development

Alan Chan, Chinasa T. Okolo|arXiv (Cornell University)|Feb 2, 2021
Ethics and Social Impacts of AI34 references10 citations
TL;DR

This paper argues that global inclusion in AI development is insufficient to address systemic inequality, as current practices—such as outsourcing data labeling and establishing foreign AI labs—reinforce power imbalances. Instead, it proposes an AI-focused import substitution industrialization (ISI) policy to enable the Global South to build domestic AI capabilities, ensuring equitable access to profits and contextual relevance in AI systems.

ABSTRACT

Those best-positioned to profit from the proliferation of artificial intelligence (AI) systems are those with the most economic power. Extant global inequality has motivated Western institutions to involve more diverse groups in the development and application of AI systems, including hiring foreign labour and establishing extra-national data centers and laboratories. However, given both the propensity of wealth to abet its own accumulation and the lack of contextual knowledge in top-down AI solutions, we argue that more focus should be placed on the redistribution of power, rather than just on including underrepresented groups. Unless more is done to ensure that opportunities to lead AI development are distributed justly, the future may hold only AI systems which are unsuited to their conditions of application, and exacerbate inequality.

Motivation & Objective

  • To examine the limitations of current inclusion practices in AI development, particularly in the Global South.
  • To highlight how top-down, Western-led AI initiatives perpetuate global inequality despite efforts to include diverse contributors.
  • To argue that inclusion alone is insufficient without redistribution of power and control over AI development.
  • To propose an AI-focused import substitution industrialization (ISI) policy as a structural solution for equitable AI development.
  • To advocate for shifting from low-productivity roles (e.g., data labeling) to high-productivity domestic AI innovation in the Global South.

Proposed method

  • Analyzing publication indices from major AI conferences (NeurIPS 2020, ICML 2020) to map global representation in AI research.
  • Evaluating the geographic and cultural bias in widely used datasets like ImageNet and OpenImages, which underrepresent Global South contexts.
  • Examining the global labor geography of data labeling, particularly the reliance on low-wage workers from sub-Saharan Africa and Southeast Asia via platforms like Amazon Mechanical Turk.
  • Drawing parallels between historical import substitution industrialization (ISI) policies and a proposed AI-focused ISI model for the Global South.
  • Proposing state-led investments in AI education, infrastructure, and domestic R&D, with controlled foreign involvement and profit reinvestment.
  • Highlighting successful regional initiatives like Deep Learning Indaba and Khipu as models for indigenous AI development.

Experimental results

Research questions

  • RQ1How do current inclusion practices in AI development reproduce global power imbalances despite increased participation from the Global South?
  • RQ2To what extent do biased datasets and outsourced data labeling undermine the contextual relevance and fairness of AI systems in the Global South?
  • RQ3What structural alternatives exist to top-down, foreign-led AI development that could enable equitable technological sovereignty in the Global South?
  • RQ4How might an AI-focused import substitution industrialization (ISI) policy enable the Global South to capture long-term economic and technical benefits from AI?
  • RQ5What role can indigenous AI education and research networks play in shifting agency from foreign institutions to local communities?

Key findings

  • Of the top 10 countries by publication index at NeurIPS 2020 and ICML 2020, none were from Latin America, Africa, or Southeast Asia, with Vietnam being the highest-ranked at 27th.
  • Eight of the top 10 institutions by publication index were based in the United States, with no universities or companies from Africa or Latin America in the top 100.
  • ImageNet and OpenImages datasets are Eurocentric and perform significantly worse on images from the Global South, such as misclassifying grooms from Ethiopia and Pakistan.
  • Over 80% of data preparation tasks in machine learning involve data collection, cleaning, and labeling, with the data labeling market projected to reach $6.5 billion by 2027.
  • Less than 2% of Amazon Mechanical Turk workers come from the Global South, with most data labeling labor outsourced to low-wage workers in sub-Saharan Africa and Southeast Asia.
  • Historical success of ISI in South Korea—through state-led industrial policy, subsidies, and education investment—offers a model for AI development in the Global South, provided similar conditions are met.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.