Challenge: Indian subcontinent is home to diverse languages written in multiple scripts . widespread use of romanization and lack of standardization means accurate transliteration models form a critical component in the NLP stack for Indian languages used by over 735 million Internet users.
Approach: They propose to build a transliteration dataset using monolingual and parallel corpora and human annotators.
Outcome: The proposed model improves accuracy by 15% on the Dakshina test set and establishes strong baselines on the Aksharantar test set.

Similar Papers

A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)

Copied to clipboard

Challenge: We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script.
Approach: They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic.
Outcome: The proposed training recipe improves multilingual transliteration for Indic languages.
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks.
Approach: They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better .
Outcome: The proposed model improves on existing models and is open-source with permissive licenses.
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages (2020.findings-emnlp)

Copied to clipboard

Challenge: In this paper, we present NLP resources for 11 major Indian languages . distributional representations are the cornerstone of modern NLP, authors say .
Approach: They introduce NLP resources for 11 major Indian languages from two major language families . monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . they also compile a benchmark for Indian language NLU to evaluate their results .
Outcome: The monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . the pre-trained language models are based on the compact ALBERT model .
IndicXNLI: Evaluating Multilingual Inference for Indian Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Indic NLP has made rapid advances in terms of corpora and pre-trained models, but benchmark datasets on standard NLU tasks are limited.
Approach: They propose to use an NLI dataset for 11 Indic languages to test their accuracy.
Outcome: The proposed dataset provides useful insights into the behaviour of pre-trained models for a diverse set of languages.
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)

Copied to clipboard

Challenge: IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages.
Approach: They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families.
Outcome: Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages.
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)

Copied to clipboard

Challenge: low-resource languages like Bangla are limited by the lack of datasets.
Approach: They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla.
Outcome: The proposed datasets show that the proposed methods can enrich romanized Bangla.
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages (2024.findings-acl)

Copied to clipboard

Challenge: Using INDICVOICES, we build the first ASR model to support all 22 languages listed in the 8th Schedule of the Constitution of India.
Approach: They propose a dataset of natural and spontaneous speech from 16237 speakers covering 145 Indian districts and 22 languages.
Outcome: The proposed dataset contains 7348 hours of read, extempore and conversational audio from 16237 speakers covering 145 Indian districts and 22 languages.
Bhasa-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages (2023.acl-short)

Copied to clipboard

Challenge: Existing tools for language identification are noisy, small and similar to high-resource languages.
Approach: They create a language identification test set for native-script and romanized text which spans all 22 Indic languages and train a model for romanized script.
Outcome: The proposed model improves on native-script and romanized script, and is competitive or better than existing LIDs.
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages (2022.tacl-1)

Copied to clipboard

Challenge: We present Samanantar, the largest publicly available parallel corpora collection for Indic languages . based on existing corporative, there has been limited benefit for resource-poor languages despite the lack of parallel corporals and monolingual corporata.
Approach: They compile 12.4 million sentence pairs from existing corpora and mine 37.4 million from the Web.
Outcome: The proposed model outperforms existing models and benchmarks on public datasets.
TRANSLIT: A Large-scale Name Transliteration Resource (2020.lrec-1)

Copied to clipboard

Challenge: Transliteration is the process of expressing a proper name from a source language in the characters of a target language.
Approach: They present a large-scale corpus of transliterated names in 180 languages . they use machine learning to train automatic transliteration .
Outcome: The proposed system achieves 92% accuracy on identification of transliterated pairs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations