Challenge: Multilingual language models are widely used to extend NLP systems to low-resource languages.
Approach: They pre-train over 10,000 monolingual and multilingual language models for over 250 languages including multiple language families that are under-studied in NLP.
Outcome: The results show that adding multilingual data improves low-resource language modeling performance, similar to increasing low-source dataset sizes by up to 33%.

Similar Papers

Mini But Mighty: Efficient Multilingual Pretraining with Linguistically-Informed Data Selection (2023.findings-eacl)

Copied to clipboard

Challenge: AfriBERTa shows that training transformer models from scratch on 1GB of data from many unrelated African languages outperforms massively multilingual models on downstream NLP tasks.
Approach: They propose that training on smaller amounts of data but from related languages could match the performance of models trained on large, unrelated data.
Outcome: The proposed model outperforms models trained on large, unrelated datasets on downstream NLP tasks.
The Less the Merrier? Investigating Language Representation in Multilingual Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Multilingual models can be used to integrate multiple languages into one model and use cross-language transfer learning to improve performance for different NLP tasks.
Approach: They propose to include languages in popular multilingual models and to use cross-language transfer learning to improve performance for different NLP tasks.
Outcome: The proposed models perform better on downstream tasks for seen and unseen languages than community-centered models for low-resource languages.
Assessing the Role of Data Quality in Training Bilingual Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that adding more languages can degrade performance for some languages while improving others.
Approach: They propose a data filtering strategy to select high-quality bilingual training data with only high quality English data.
Outcome: The proposed approach improves bilingual model performance by 2–4% and reduces bilingual models performance gaps to 1%.
Targeted Multilingual Adaptation for Low-resource Language Families (2024.findings-emnlp)

Copied to clipboard

Challenge: Massively multilingual models are known to have limited utility in any one language, and to perform poorly on low-resource languages.
Approach: They propose to adapt a pre-trained multilingual model to a language family and evaluate its performance on two downstream tasks and 11 evaluation languages.
Outcome: The proposed model outperforms mono- and multilingual models on two downstream tasks and 11 evaluation languages.
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing practices of fine-tuning and evaluating multilingual large language models may not align with this objective due to a heavy reliance on translation.
Approach: They propose to use translated or native instruction data to fine-tune multilingual large language models.
Outcome: The proposed model can be fine tuned and evaluated in multilingual large language models . the results show that native or translated data can be used to compare model performance .
Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models (2022.emnlp-main)

Copied to clipboard

Challenge: a large number of pretraining corpora are not publicly available, and it is unclear how much foreign language data exists in monolingual models.
Approach: They propose to use English pretraining corpora to analyze their language composition . they find that even when less than 1% of data is not English, it facilitates cross-lingual transfer .
Outcome: The proposed model is not truly monolingual when pretrained at scale, the authors show . they show that even when less than 1% of data is not English, it facilitates cross-lingual transfer .
How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM (2025.coling-main)

Copied to clipboard

Challenge: Many large language models (LLMs) support many languages, while others only support a few, e.g. the Llama series.
Approach: They present a case study on BLOOM to understand three pertinent factors affecting performance: the number of languages, language exposure, and similarity between training and test languages.
Outcome: The proposed model can be used to perform multilingual tasks on 1 to 52 languages.
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Multilingual language models often underperform monolingual ones due to inter-language competition for model parameters.
Approach: They propose Cross-lingual Expert Language Models (X-ELM) which mitigates inter-language competition by independently training language models on subsets of the multilingual corpus.
Outcome: The proposed model outperforms jointly trained multilingual models across all 16 considered languages and transfer the gains to downstream tasks.
Multilingual BERT has an accent: Evaluating English influences on fluency in multilingual models (2023.findings-eacl)

Copied to clipboard

Challenge: Multilingual models can improve NLP performance on low-resource languages by leveraging higher-resourced languages, but they also reduce average performance on all languages.
Approach: They propose a method to evaluate multilingual models by asking if models predict languages with an 'English accent' they propose to use grammatical structure bias to determine if multilingual model is biased toward English-like setting .
Outcome: The proposed method compares the fluency of multilingual models to the fluencies of monolingual Spanish and Greek models.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations