Challenge: Continued pretraining (CPT) is a practical route to language adaptation, but improvements on demanding capabilities such as mathematical reasoning are limited.
Approach: They propose to use CPT to adapt large language models to African languages . they use math, code, and synthetic translated data to analyze their models .
Outcome: The proposed models improve on multilingual benchmarks and document-level translation.

Similar Papers

AfriInstruct: Instruction Tuning of African Languages for Diverse Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) for African languages perform worse compared to high-resource languages.
Approach: They propose a model that specializes in instruction-tuning of multiple African languages covering various tasks.
Outcome: The proposed model outperforms GPT-3.5-Turbo and other models of similar size in multiple tasks.
Emergent Abilities of Large Language Models under Continued Pre-training for Language Adaptation (2025.acl-long)

Copied to clipboard

Challenge: Existing large language models are notoriously English-centric, and their performance has been reported to drop significantly in lessresourced languages.
Approach: They propose a language-agnostic benchmark for in-context learning that reveals catastrophic forgetting early on CPT when English is not included.
Outcome: The proposed method does not impact validation perplexity but is critical for emergence of downstream capabilities in the target language.
Towards Effective and Efficient Continual Pre-training of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks.
Approach: They propose a Continual pre-training method that can greatly improve Chinese language ability and scientific reasoning ability of LLMs.
Outcome: The proposed method can greatly improve Chinese language ability and scientific reasoning ability of LLMs.
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made significant strides towards Artificial General Intelligence, but training them from scratch is prohibitively expensive.
Approach: They propose to continuously pre-train LLMs from existing pre-trained LLM models by using a set of parameters instead of randomly initializing them.
Outcome: The proposed approach saves significant resources and accelerates convergence and performance.
AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in multimodal and speech-native large language models have delivered impressive speech recognition, translation, understanding, and question-answering capabilities for high-resource languages.
Approach: They propose to benchmark African languages and African-accented French, Arabic, and 100+ African English accents across 20 African languages.
Outcome: The proposed model outperforms traditional speech transcription and translation models in African languages and non-native French or English accents.
Efficient Continual Pre-training of LLMs for Low-resource Languages (2025.naacl-industry)

Copied to clipboard

Challenge: Open-source large language models (LLMs) are a promising tool for low-resource languages . however, there is still a substantial performance gap between high-resourced languages and LRLs .
Approach: They develop an algorithm to select a subset of texts from a larger corpus and use it to select tokens for LLMs.
Outcome: The proposed algorithm reduces the cost of continual pre-training (CPT) with large amounts of language-specific data.
Continued Pretraining and Interpretability-Based Evaluation for Low-Resource Languages: A Galician Case Study (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to remarkable improvements in language understanding and text generation.
Approach: They propose a framework to evaluate large language models for underrepresented languages . they examine CPT strategies for languages with limited representation in multilingual models .
Outcome: The proposed evaluation framework is based on the case of Galician language . it assesses trade-offs between linguistic enrichment and task-solving capabilities .
Multilingual Language Model Pretraining using Machine-translated Data (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for collecting and filtering multilingual web data lead to most languages lagging behind English performance due to the Internet's English-centric nature.
Approach: They propose to translate a high-quality English web corpus into nine languages and pretrain a 1.3B-parameter model on it.
Outcome: The proposed model matches or outperforms multilingual LLMs of similar size across Non-English understanding and reasoning tasks despite being trained on an order of magnitude less data.
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM𝛥 Integration into Upcycled MoE (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are expensive and require extensive Continued Pre-Training and data-intensive alignment to expand.
Approach: They propose a method which upcycles a dense model into a Mixture-of-Experts architecture, allocating different experts to different languages.
Outcome: Experiments show that the proposed model upcycles a dense model into a Mixture-of-Experts(MoE) architecture, allocating different experts to different languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations