Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs (2025.coling-main)
Copied to clipboard
Tamzeed Mahfuz, Satak Kumar Dey, Ruwad Naswan, Hasnaen Adil, Khondker Salman Sayeed, Haz Sameen Shahgir
| Challenge: | a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages. |
| Approach: | They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data . |
| Outcome: | The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks . |
Similar Papers
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)
Copied to clipboard
| Challenge: | Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
| Approach: | They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages. |
| Outcome: | The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting . |
| Approach: | They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models . |
| Outcome: | The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects. |
TigerLLM - A Family of Bangla Large Language Models (2025.acl-short)
Copied to clipboard
| Challenge: | linguistic disparity is particularly evident for Bangla, the 5th most spoken language . open-source Bangla LLMs have limited reproducibility and performance gaps . |
| Approach: | They propose a family of Bangla LLMs that outperform open-source alternatives and benchmarks and establish a new benchmark for future Bangla language modeling. |
| Outcome: | The proposed models outperform existing models and outperformed proprietary models across six benchmarks. |
BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP (2024.lrec-main)
Copied to clipboard
Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, Enamul Hoque
| Challenge: | Large Language Models (LLMs) have emerged as one of the most important breakthroughs in natural language processing. |
| Approach: | They propose to evaluate LLMs in Bengali to benchmark their performance . they select Bangla NLP tasks such as text summarization, question answering, paraphrasing . |
| Outcome: | The proposed model performs better in some tasks than current models, but in most tasks, it is poor . |
It’s All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Low-resource languages, especially those written in rare scripts, remain unsupported by large language models due to lack of training data. |
| Approach: | They evaluate 20 under-represented languages across three state-of-the-art multilingual LLMs and compare their methods to parameter-efficient fine-tuning. |
| Outcome: | The proposed methods compare with parameter-efficient fine-tuning (PEFT) on low-resource languages. |
LLMs for Extremely Low-Resource Finno-Ugric Languages (2025.findings-naacl)
Copied to clipboard
| Challenge: | Low-resource languages such as those in the Finno-Ugric family are underrepresented in large language models. |
| Approach: | They propose to develop large language models for extremely low-resource languages . they focus on Vro, Livonian, and Komi, which are underrepresented . |
| Outcome: | The proposed models cover almost the entire cycle of creation, from data collection to instruction tuning and evaluation. |
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. |
| Approach: | They evaluate the performance of large language models and their generation strategies in 11 different languages using 3 NLP tasks and 4 open-source LLMs. |
| Outcome: | The proposed generation strategies and their combinations yield strong results across 11 languages, including several extremely low-resource ones. |
Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech Identification (2025.emnlp-main)
Copied to clipboard
| Challenge: | despite Bengali being among the most spoken languages, the NLP efforts on it remain limited. |
| Approach: | They present a dataset that includes Bengali figures of speech on six poets . they deploy state-of-the-art Large Language Models to the dataset and fine-tune the best models . |
| Outcome: | The proposed dataset reveals that two open-source LLMs perform better than others in Bengali . the framework can be reproduced for English and other low-resource languages . |
Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a survey examines the current efforts and challenges of NLP models for South Asian languages . there are more than 650 languages in South Asia, but many have very limited computational resources or are missing from existing models. |
| Approach: | a survey examines efforts and challenges of NLP for South Asian languages . they focus on transformer-based models such as BERT, T5, & GPT . findings highlight substantial issues, including missing data in critical domains . |
| Outcome: | The findings highlight significant issues, including missing data in critical domains . the survey aims to raise awareness within the NLP community for more targeted data curation . |
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment. |
| Approach: | They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge. |
| Outcome: | The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses. |