Challenge: Low-resource languages, especially those written in rare scripts, remain unsupported by large language models due to lack of training data.
Approach: They evaluate 20 under-represented languages across three state-of-the-art multilingual LLMs and compare their methods to parameter-efficient fine-tuning.
Outcome: The proposed methods compare with parameter-efficient fine-tuning (PEFT) on low-resource languages.

Similar Papers

LLMs Are Few-Shot In-Context Low-Resource Language Learners (2024.naacl-long)

Copied to clipboard

Challenge: In-context learning (ICL) empowers large language models to perform diverse tasks in underrepresented languages using only short in-contrast information.
Approach: They extensively assess the effectiveness of in-context learning with LLMs in low-resource languages . they also identify the shortcomings of in context label alignment .
Outcome: The proposed approach improves understanding quality of low-resource languages by closing the language gap in the target language.
Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks.
Approach: They propose to make Large Language Models (LLMs) operating in 0-shot or few-shot settings as efficient as 0- shot text classifiers by leveraging a small number of samples.
Outcome: The proposed model is able to perform better on multiple datasets than existing models on 0-shot or few-shot settings.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.
Can Large Language Models Translate Unseen Languages in Underrepresented Scripts? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive performance in machine translation, but struggle with unseen low-resource languages.
Approach: They propose a benchmark to evaluate translation for Mongolian and Yi using linguistic resources.
Outcome: The proposed model can translate Mongolian (in traditional script) and Yi with the help of linguistic resources, but is limited in its ability to handle these languages effectively.
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
Approach: They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages.
Outcome: The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
Multimodal In-context Learning for ASR of Low-resource Languages (2026.findings-acl)

Copied to clipboard

Challenge: In-context learning with large language models addresses this limitation, but prior work focuses on high-resource languages covered during training and text-only settings.
Approach: They propose to use multimodal ICL to learn unseen languages with multimodal learning to improve ASR in large language models.
Outcome: The proposed model outperforms existing models on unseen languages with multimodal ICL (MICL) and cross-lingual transfer learning matches or outperformed models without using target-language data.
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective (2026.acl-long)

Copied to clipboard

Challenge: Prior studies comparing FT and ICL have yielded mixed and inconclusive results due to inconsistent experimental setups.
Approach: They propose a formal language learning task with precise language boundaries, controlled string sampling, and no data contamination to enable a rigorous comparison.
Outcome: The proposed task offers precise language boundaries, controlled string sampling, and no data contamination.
Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment (2023.acl-long)

Copied to clipboard

Challenge: a handful of studies have explored ICL in a cross-lingual setting . emergence of large-scale, pretrained, Transformer-based language models has marked the commencement of an avant-garde era in NLP.
Approach: They propose a novel prompt construction strategy to bridge the gap between ICL and cross-lingual text classification.
Outcome: The proposed approach outperforms random prompt selection by a large margin across three tasks using 44 different cross-lingual pairs.
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs (2025.coling-main)

Copied to clipboard

Challenge: a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages.
Approach: They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data .
Outcome: The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks .
VEEF-Multi-LLM: Effective Vocabulary Expansion and Parameter Efficient Finetuning Towards Multilingual Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a significant disadvantage for low-resource languages . VEEF-Multi-LLM-8B excels in multilingual instruction-following tasks .
Approach: They propose a low-resource multilingual large language model that expands the vocabulary for multilingual support.
Outcome: The proposed model outperforms existing models in multilingual instruction-following tasks, but lags behind English-centric models in some tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations