Challenge: 'BanglaNLG' is a comprehensive benchmark for evaluating natural language generation models in Bangla, a widely spoken yet low-resource language.
Approach: They propose to aggregate six conditional text generation tasks under the BanglaNLG benchmark and introduce a new dataset on dialogue generation in the process.
Outcome: The proposed model outperforms several multilingual models by 9% absolute gain and 32% relative gain on all of these tasks.

Similar Papers

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla (2022.findings-naacl)

Copied to clipboard

Challenge: Bangla is a widely spoken yet low-resource language in the NLP literature.
Approach: They propose a BERT-based natural language understanding model pretrainable in Bangla, a widely spoken yet low-resource language in the NLP literature.
Outcome: The proposed model outperforms multilingual and monolingual models on four NLU tasks covering text classification, sequence labeling, and span prediction.
IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks .
Approach: They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models .
Outcome: The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages .
BanglaByT5: Byte-Level Modelling for Bangla (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have achievedremarkable success across various natural lan-guage processing tasks.
Approach: They propose a byte-level encoder-decoder model specifically tailored for Bangla.
Outcome: The proposed model outperforms existing models in gen-erative and classification tasks and surpasses several multilingual and larger models.
BanglaParaphrase: A High-Quality Bangla Paraphrase Dataset (2022.aacl-short)

Copied to clipboard

Challenge: Bangla is considered a low resource language in terms of language processing.
Approach: They propose a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline.
Outcome: The proposed pipeline ensures quality by preserving both semantics and diversity, making it particularly useful to enhance other Bangla datasets.
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)

Copied to clipboard

Challenge: low-resource languages like Bangla are limited by the lack of datasets.
Approach: They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla.
Outcome: The proposed datasets show that the proposed methods can enrich romanized Bangla.
BanglaSTEM: A Parallel Corpus and Term-Weighted Evaluation for Technical Bangla-English Translation (2026.acl-srw)

Copied to clipboard

Challenge: Large language models excel at technical problem solving in English but struggle when questions are posed in Bangla.
Approach: They propose a dataset of 5,000 Bangla-English sentence pairs to align technical terms . they use OCR to extract matching passages from bilingual textbooks .
Outcome: The proposed pipeline extracts matching passages from bilingual textbooks and uses them to align sentences and mark technical terms.
TigerLLM - A Family of Bangla Large Language Models (2025.acl-short)

Copied to clipboard

Challenge: linguistic disparity is particularly evident for Bangla, the 5th most spoken language . open-source Bangla LLMs have limited reproducibility and performance gaps .
Approach: They propose a family of Bangla LLMs that outperform open-source alternatives and benchmarks and establish a new benchmark for future Bangla language modeling.
Outcome: The proposed models outperform existing models and outperformed proprietary models across six benchmarks.
BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have emerged as one of the most important breakthroughs in natural language processing.
Approach: They propose to evaluate LLMs in Bengali to benchmark their performance . they select Bangla NLP tasks such as text summarization, question answering, paraphrasing .
Outcome: The proposed model performs better in some tasks than current models, but in most tasks, it is poor .
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarking datasets for Bangla LLMs are not available for all languages.
Approach: They present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes.
Outcome: The proposed model outperforms existing models in Bangla, but not always in the first place.
BanglaRQA: A Benchmark Dataset for Under-resourced Bangla Language Reading Comprehension-based Question Answering with Diverse Question-Answer Types (2022.findings-emnlp)

Copied to clipboard

Challenge: a lack of diverse and comprehensive question-answering datasets exists in under-resourced languages like Bangla.
Approach: They propose a reading comprehension-based Bangla question-answering dataset . the dataset includes answerable and unanswerable questions covering four categories of questions .
Outcome: The proposed dataset shows that it performs well as a training resource in high-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations