Challenge: In this paper, we present NLP resources for 11 major Indian languages . distributional representations are the cornerstone of modern NLP, authors say .
Approach: They introduce NLP resources for 11 major Indian languages from two major language families . monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . they also compile a benchmark for Indian language NLU to evaluate their results .
Outcome: The monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . the pre-trained language models are based on the compact ALBERT model .

Similar Papers

IndicXNLI: Evaluating Multilingual Inference for Indian Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Indic NLP has made rapid advances in terms of corpora and pre-trained models, but benchmark datasets on standard NLU tasks are limited.
Approach: They propose to use an NLI dataset for 11 Indic languages to test their accuracy.
Outcome: The proposed dataset provides useful insights into the behaviour of pre-trained models for a diverse set of languages.
Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in Natural Language Understanding are driven by pretrained multilingual models, which can potentially reduce the performance gap between high-resource languages through zero-shot knowledge transfer.
Approach: They propose to create a human-supervised benchmark for Indic languages, IndicXTREME, with nine diverse NLU tasks covering 20 languages.
Outcome: The proposed model improves on the monolingual corpora, IndicCorp, and IndicBERT in Indic languages with 105 evaluation sets across languages and tasks.
IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages (2022.emnlp-main)

Copied to clipboard

Challenge: IndicNLG is a non-English language that is hampered by the scarcity of datasets.
Approach: They propose to create a dataset for natural language generation for 11 Indic languages . they use a set of pre-trained models to train multilingual models .
Outcome: The proposed datasets show that pre-trained models perform well in multilingual and monolingual tasks.
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)

Copied to clipboard

Challenge: IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages.
Approach: They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families.
Outcome: Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages.
MILU: A Multi-task Indic Language Understanding Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on English, leaving substantial gaps in assessing LLM capabilities in low-resource and linguistically diverse languages.
Approach: They propose a multi-task indic language understanding benchmark to assess LLMs in low-resource languages.
Outcome: The new benchmark spans 8 domains and 41 subjects across 11 Indic languages, reflecting general and culturally specific knowledge.
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources (2026.acl-long)

Copied to clipboard

Challenge: Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties.
Approach: They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
Outcome: The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)

Copied to clipboard

Challenge: We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script.
Approach: They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic.
Outcome: The proposed training recipe improves multilingual transliteration for Indic languages.
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages (2022.tacl-1)

Copied to clipboard

Challenge: We present Samanantar, the largest publicly available parallel corpora collection for Indic languages . based on existing corporative, there has been limited benefit for resource-poor languages despite the lack of parallel corporals and monolingual corporata.
Approach: They compile 12.4 million sentence pairs from existing corpora and mine 37.4 million from the Web.
Outcome: The proposed model outperforms existing models and benchmarks on public datasets.
Naamapadam: A Large-Scale Named Entity Annotated Data for Indic Languages (2023.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a fundamental task in natural language processing (NLP).
Approach: They present the largest publicly available Named Entity Recognition dataset for the 11 major Indian languages from two language families.
Outcome: The proposed dataset is the largest publicly available Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families.
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights (2026.findings-acl)

Copied to clipboard

Challenge: Existing tokenizers are often skewed towards high-resource languages limiting their effectiveness for linguistically diverse and morphologically rich languages.
Approach: They evaluate multilingual tokenization across 17 Indic languages spanning 11 scripts and two language families.
Outcome: The proposed method improves tokenization quality and vocabulary size in 17 languages . poor tokenization can lead to increase in sequence lengths, fragment meaningful units, weaken model's ability to capture linguistic structure and semantics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations