Skip to main content
QUICK REVIEW

[Paper Review] A Multilingual African Embedding for FAQ Chatbots

Aymen Ben Elhaj Mabrouk, Moez Ben HajHmida|arXiv (Cornell University)|Mar 16, 2021
AI in Service Interactions8 references4 citations
TL;DR

This paper introduces a multilingual, multilingual-dialectal FAQ chatbot for African crisis communication, using a modified StarSpace embedding tailored for African languages including Arabic, French, English, Tunisian, Igbo, Yorùbá, and Hausa. The system achieves high user engagement and satisfactory response quality, with 7.4 average messages per conversation and a 68% sensibleness and 60% specificity score in human evaluation, demonstrating effectiveness in multilingual, low-resource African contexts.

ABSTRACT

Searching for an available, reliable, official, and understandable information is not a trivial task due to scattered information across the internet, and the availability lack of governmental communication channels communicating with African dialects and languages. In this paper, we introduce an Artificial Intelligence Powered chatbot for crisis communication that would be omnichannel, multilingual and multi dialectal. We present our work on modified StarSpace embedding tailored for African dialects for the question-answering task along with the architecture of the proposed chatbot system and a description of the different layers. English, French, Arabic, Tunisian, Igbo,Yorùbá, and Hausa are used as languages and dialects. Quantitative and qualitative evaluation results are obtained for our real deployed Covid-19 chatbot. Results show that users are satisfied and the conversation with the chatbot is meeting customer needs.

Motivation & Objective

  • To address the lack of reliable, official, and accessible multilingual crisis communication in African local dialects during the Covid-19 pandemic.
  • To develop an omnichannel, multilingual, and multi-dialectal chatbot that serves citizens in both official languages and local dialects such as Tunisian Arabic, Igbo, Yorùbá, and Hausa.
  • To overcome the absence of pretrained language models and domain-specific embeddings for African dialects by creating a custom embedding model.
  • To improve access to reliable information for vulnerable populations, including rural and illiterate citizens, by digitizing institutional services.
  • To evaluate the chatbot’s performance quantitatively and qualitatively, ensuring user satisfaction and response relevance.

Proposed method

  • Collected official, reliable Covid-19 Q&A data from Tunisian and Nigerian governmental and NGO sources.
  • Augmented the dataset with chitchat-like interactions (e.g., greetings, jokes) to improve conversational fluency.
  • Split the dataset into two categories: Frequently Asked Questions (FAQ) and chitchat for intent classification.
  • Modified the StarSpace embedding model to support multilingual and multi-dialectal representation learning for African languages and dialects.
  • Designed a chatbot architecture capable of handling multiple languages and dialects with zero-shot or few-shot adaptation via the custom embedding layer.
  • Deployed the chatbot across multiple channels (e.g., Messenger) and collected real-world interaction logs for evaluation.

Experimental results

Research questions

  • RQ1Can a multilingual, multi-dialectal chatbot effectively serve African citizens in both official languages and local dialects during a crisis?
  • RQ2How does a modified StarSpace embedding model perform in learning semantic representations for low-resource African dialects?
  • RQ3What is the real-world user engagement and satisfaction level of a crisis chatbot deployed in African multilingual contexts?
  • RQ4To what extent can the chatbot provide sensible and specific responses across diverse African languages and dialects?
  • RQ5Can such a system improve access to reliable information for vulnerable populations, including rural and illiterate users?

Key findings

  • The chatbot achieved a high interaction rate, with 8,153 unique users on Messenger between March 23 and June 18, 2020, and an average of 7.4 questions per conversation.
  • The stickiness rate (daily active users divided by monthly active users) reached 16.85%, indicating sustained user engagement.
  • Human evaluation using the Sensibleness and Specificity Average (SSA) metric showed 68% sensibleness and 60% specificity, indicating high-quality, contextually relevant responses.
  • The user base was diverse, with significant engagement from age groups 18–34, and notable representation from both females and males across age groups.
  • The chatbot successfully handled questions in multiple languages including English, French, Arabic, Tunisian, Igbo, Yorùbá, and Hausa without predefined scenarios.
  • The modified StarSpace embedding model enabled effective cross-lingual and cross-dialectal understanding, supporting deployment across multiple African languages and dialects.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.