Pula: Training Large Language Models for Setswana (2025.naacl-long)

Copied to clipboard

Challenge: Setswana is a Bantu language spoken by an estimated five to ten million people worldwide.
Approach: They propose to make setswana-based models available for the first time using data available from setswa and setswegian databases.
Outcome: The proposed models outperform GPT-4o and Gemini 1.5 Pro on English-Setswana translation tasks and achieve state-of-the-art performance on Setswanan reasoning tasks.

Similar Papers

NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages (2024.lrec-main)

Copied to clipboard

Challenge: Nguni languages have over 20 million home language speakers in South Africa . there has been considerable growth in the datasets for these languages, but no analysis of the performance of NLP models for these language has been reported across languages and tasks.
Approach: They compile publicly available datasets for natural language understanding and generation, spanning 6 tasks and 11 datasets.
Outcome: The proposed models outperform existing models and large-scale adapted models on cross-lingual transfer and machine translation.
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data (2024.emnlp-industry)

Copied to clipboard

Challenge: Large language models exhibit significant performance discrepancies between high- and low-resource languages.
Approach: They present an open-source multilingual LLM with 8 billion parameters and a multilingual instruction dataset.
Outcome: The proposed model achieves consistent multilingual representations across languages.
BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification (2022.lrec-1)

Copied to clipboard

Challenge: Large-scale monolingual pre-trained language models have shown promising results for high-resource as well as lowresource languages, especially for text classification.
Approach: They provide a set of recommendations for using pre-trained models for Sinhala text classification and introduce new annotated datasets useful for future research.
Outcome: The proposed models are far superior to existing models for Sinhala and set a strong baseline for text classification when fine-tuned.
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages (2026.acl-long)

Copied to clipboard

Challenge: Multilingual large language models are expensive to pretrain and suffer from imbalances across languages and datasets.
Approach: They propose a family of Indian language-only autoregressive language models trained on open-source language-specific data for the five most spoken Indian languages.
Outcome: The proposed model outperforms most larger models up to 8B across all five languages.
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown continuously improving multilingual capabilities.
Approach: They evaluate the ability of open LLMs to handle multilingual machine translation tasks using a parallel-first monolingual-second data mixing strategy.
Outcome: The proposed model outperforms state-of-the-art models and achieves competitive performance with Google Translate and GPT-4-turbo.
Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages (2023.acl-long)

Copied to clipboard

Challenge: Lack of LLMs supporting low-resource languages is a serious impediment to bringing NLP to all of the world.
Approach: They create a model that scales LLMs horizontally and a corpus that covers 511 low-resource languages.
Outcome: The proposed model improves on five diverse tasks across low- and high-resource languages.
Sinhala Encoder-only Language Models and Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language models (LMs) have produced excellent results in many NLP tasks, but their effectiveness is highly dependent on available pre-training resources.
Approach: They propose to collect the largest monolingual corpus for Sinhala and compile a benchmark and evaluate LMs on it.
Outcome: The proposed language models outperform the popular multilingual LMs in downstream NLP tasks.
XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large multilingual models rely on a single vocabulary shared across 100+ languages . this vocabulary bottleneck limits the representational capabilities of multilingual model XLM-R .
Approach: They propose a new approach for scaling to large multilingual vocabularies by de-emphasizing token sharing between languages with little lexical overlap and assigning vocabulary capacity to achieve sufficient coverage for each individual language.
Outcome: The proposed model outperforms XLM-R on all language tasks and is particularly effective on low-resource tasks.
LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit exceptional translation capabilities in high-resource language tasks, yet their effectiveness in low-resourced languages is suboptimal.
Approach: They conduct extensive multilingual continual pre-training on the LLaMA series models and develop LLiMAX for translation support across more than 100 languages.
Outcome: The proposed model achieves higher translation performance than existing open-source models and performs on-par with specialized translation model on the Flores-101 benchmark.
AfriInstruct: Instruction Tuning of African Languages for Diverse Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) for African languages perform worse compared to high-resource languages.
Approach: They propose a model that specializes in instruction-tuning of multiple African languages covering various tasks.
Outcome: The proposed model outperforms GPT-3.5-Turbo and other models of similar size in multiple tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations