[Paper Review] Global Readiness of Language Technology for Healthcare: What would it Take to Combat the Next Pandemic?
This paper evaluates the global readiness of language technology (LT) for healthcare, focusing on chatbot development for pandemic response across 15 Asian and African languages. Using intent classification with multilingual models and commercial frameworks, it finds a 20–30% performance drop in low-resource languages (classes 0–2), revealing stark disparities in LT accessibility and proposing targeted investment in 'bridge' languages and domain-adapted NLP systems to ensure global pandemic preparedness.
The COVID-19 pandemic has brought out both the best and worst of language technology (LT). On one hand, conversational agents for information dissemination and basic diagnosis have seen widespread use, and arguably, had an important role in combating the pandemic. On the other hand, it has also become clear that such technologies are readily available for a handful of languages, and the vast majority of the global south is completely bereft of these benefits. What is the state of LT, especially conversational agents, for healthcare across the world's languages? And, what would it take to ensure global readiness of LT before the next pandemic? In this paper, we try to answer these questions through survey of existing literature and resources, as well as through a rapid chatbot building exercise for 15 Asian and African languages with varying amount of resource-availability. The study confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help us prioritize research and investment strategies in LT for healthcare.
Motivation & Objective
- To assess the current state of language technology (LT) readiness for healthcare chatbots across the world’s languages, especially in low-resource settings.
- To investigate the feasibility of building practical, multilingual COVID-19 FAQ chatbots for under-resourced Asian and African languages using existing commercial and open-source tools.
- To identify key gaps in language resource availability and model performance that hinder equitable pandemic response technologies.
- To develop a global LT readiness map for pandemic response, using country-level language demographics and model performance scores.
- To recommend targeted research and investment strategies to ensure global LT readiness before the next pandemic.
Proposed method
- Selected 15 Asian and African languages across six resource classes (0–5), based on linguistic resource availability, to evaluate LT performance.
- Built intent classifiers for each language using multilingual transformer models (mBERT and XLM-R) and commercial frameworks (Google Dialogflow, Microsoft Bot Framework).
- Evaluated performance on a standardized COVID-19 FAQ intent classification task, measuring F1-scores relative to English baseline.
- Conducted entity recognition experiments on a subset of languages to assess named entity detection capabilities in low-resource settings.
- Calculated language-level readiness scores ($ r_l $) based on model performance, then aggregated to country-level readiness ($ r_c $) using population-weighted language distributions.
- Applied Jenks’ natural breaks optimization to classify countries into five pandemic-readiness levels (Extremely ill-prepared to Fully prepared), generating a global heatmap.
Experimental results
Research questions
- RQ1Which languages currently have sufficient language technology infrastructure to support practical healthcare chatbots for pandemic response?
- RQ2How does performance of multilingual NLU models degrade in low-resource languages (classes 0–2) compared to high-resource languages like English?
- RQ3To what extent do commercial chatbot frameworks and multilingual models support under-resourced languages in Africa and Asia?
- RQ4What are the key linguistic and technical barriers preventing effective deployment of healthcare chatbots in low-resource languages?
- RQ5How can global LT readiness for healthcare be improved before the next pandemic through strategic research and investment?
Key findings
- A 20–30% drop in intent classification F1-score was observed for class 0–2 languages (e.g., Hausa, Somali, Marathi, Sinhala) compared to English, even when using state-of-the-art multilingual models and commercial frameworks.
- Despite using mBERT and XLM-R, performance for African languages and some Indian languages (e.g., Marathi, Sinhala) remained significantly below English baseline, indicating persistent resource gaps.
- Only a few European and Asian languages (e.g., French, Chinese, Korean, Bengali) beyond English reached high readiness levels, while widely spoken languages like Swahili and Hausa remain under-resourced.
- Countries in sub-Saharan Africa, South Asia, and parts of Latin America (e.g., Zambia, Bolivia, Guatemala) were classified as 'Extremely ill-prepared' due to high proportions of low-resource languages.
- The presence of a geographically or linguistically close 'bridge' language (e.g., Hindi for Marathi) improved performance in some low-resource languages, suggesting transfer learning potential.
- Commercial chatbot frameworks and MT systems showed brittleness with domain-specific terms (e.g., 'incubation', 'COVID'), highlighting the need for domain adaptation and term injection techniques.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.