[Paper Review] A Primer on Pretrained Multilingual Language Models
A survey of multilingual language models (MLLMs) covering their architectures, objectives, data, benchmarks, and their comparative performance with monolingual models, along with future directions.
Multilingual Language Models (\MLLMs) such as mBERT, XLM, XLM-R, extit{etc.} have emerged as a viable option for bringing the power of pretraining to a large number of languages. Given their success in zero-shot transfer learning, there has emerged a large body of work in (i) building bigger \MLLMs~covering a large number of languages (ii) creating exhaustive benchmarks covering a wider variety of tasks and languages for evaluating \MLLMs~ (iii) analysing the performance of \MLLMs~on monolingual, zero-shot cross-lingual and bilingual tasks (iv) understanding the universal language patterns (if any) learnt by \MLLMs~ and (v) augmenting the (often) limited capacity of \MLLMs~ to improve their performance on seen or even unseen languages. In this survey, we review the existing literature covering the above broad areas of research pertaining to \MLLMs. Based on our survey, we recommend some promising directions of future research.
Motivation & Objective
- Review how MLLMs are built and how they differ across models.
- Summarize training objectives (monolingual and parallel) and data sources used for pretraining.
- Survey benchmarks used to evaluate MLLMs across languages and tasks.
- Discuss whether MLLMs outperform monolingual LMs for specific languages and tasks.
- Highlight methods to extend MLLMs to unseen languages and recommend future directions.
Proposed method
- Catalog architectures of representative MLLMs and their configurations (N, k, d, parameters).
- Compare pretraining objectives: MLM, CLM, MRTD, TLM, CAMLM, CLMLM, XLCO, HICTL, CLSA and related methods.
- Describe monolingual vs. parallel objective combinations and their data requirements.
- Explain pretraining data choices and language sampling via exponential smoothing to balance language exposure.
- Summarize evaluation benchmarks like XGLUE, XTREME, XTREME-R and XGLUE across tasks.
- Synthesize findings on cross-lingual transfer, bilingual tasks, and universal-pattern hypotheses.
Experimental results
Research questions
- RQ1How are different MLLMs built and how do they differ from each other?
- RQ2What benchmarks are used to evaluate MLLMs?
- RQ3Are MLLMs better than monolingual LMs for a given language?
- RQ4Do MLLMs enable zero-shot cross-lingual transfer and bilingual tasks?
- RQ5Do MLLMs reveal universal language patterns and how to extend them to new languages?
Key findings
- MLLMs vary in architecture, pretraining data, languages covered, and vocabulary size; higher capacity models generally perform better in benchmarks.
- Cross-lingual transfer ability depends on factors like shared vocabulary, representation alignment, and data size; no single model outperforms monolingual LMs consistently across all settings.
- Benchmarks such as XGLUE, XTREME, and XTREME-R are used to assess cross-lingual performance across classification, structure prediction, QA, and retrieval tasks.
- There is evidence that zero-shot cross-lingual transfer and bilingual tasks benefit from MLLMs, especially in low-resource languages, but universal interlingua patterns are not yet established.
- Extending MLLMs to unseen languages can be approached via fine-tuning, adapters, and other extension methods.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.