Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users (2023.findings-emnlp)
Copied to clipboard
Yash Madhani, Sushane Parthan, Priyanka Bedekar, Gokul Nc, Ruchi Khapra, Anoop Kunchukuttan, Pratyush Kumar, Mitesh Khapra
| Challenge: | Indian subcontinent is home to diverse languages written in multiple scripts . widespread use of romanization and lack of standardization means accurate transliteration models form a critical component in the NLP stack for Indian languages used by over 735 million Internet users. |
| Approach: | They propose to build a transliteration dataset using monolingual and parallel corpora and human annotators. |
| Outcome: | The proposed model improves accuracy by 15% on the Dakshina test set and establishes strong baselines on the Aksharantar test set. |
Similar Papers
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)
Copied to clipboard
| Challenge: | We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script. |
| Approach: | They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic. |
| Outcome: | The proposed training recipe improves multilingual transliteration for Indic languages. |
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)
Copied to clipboard
Ashwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M Khapra, Raj Dabre
| Challenge: | Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks. |
| Approach: | They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better . |
| Outcome: | The proposed model improves on existing models and is open-source with permissive licenses. |
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages (2020.findings-emnlp)
Copied to clipboard
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, Pratyush Kumar
| Challenge: | In this paper, we present NLP resources for 11 major Indian languages . distributional representations are the cornerstone of modern NLP, authors say . |
| Approach: | They introduce NLP resources for 11 major Indian languages from two major language families . monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . they also compile a benchmark for Indian language NLU to evaluate their results . |
| Outcome: | The monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . the pre-trained language models are based on the compact ALBERT model . |
IndicXNLI: Evaluating Multilingual Inference for Indian Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | Indic NLP has made rapid advances in terms of corpora and pre-trained models, but benchmark datasets on standard NLU tasks are limited. |
| Approach: | They propose to use an NLI dataset for 11 Indic languages to test their accuracy. |
| Outcome: | The proposed dataset provides useful insights into the behaviour of pre-trained models for a diverse set of languages. |
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)
Copied to clipboard
| Challenge: | low-resource languages like Bangla are limited by the lack of datasets. |
| Approach: | They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla. |
| Outcome: | The proposed datasets show that the proposed methods can enrich romanized Bangla. |
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages (2024.findings-acl)
Copied to clipboard
Tahir Javed, Janki Nawale, Eldho George, Sakshi Joshi, Kaushal Bhogale, Deovrat Mehendale, Ishvinder Sethi, Aparna Ananthanarayanan, Hafsah Faquih, Pratiti Palit, Sneha Ravishankar, Saranya Sukumaran, Tripura Panchagnula, Sunjay Murali, Kunal Gandhi, Ambujavalli R, Manickam M, C Vaijayanthi, Krishnan Karunganni, Pratyush Kumar, Mitesh Khapra
| Challenge: | Using INDICVOICES, we build the first ASR model to support all 22 languages listed in the 8th Schedule of the Constitution of India. |
| Approach: | They propose a dataset of natural and spontaneous speech from 16237 speakers covering 145 Indian districts and 22 languages. |
| Outcome: | The proposed dataset contains 7348 hours of read, extempore and conversational audio from 16237 speakers covering 145 Indian districts and 22 languages. |
Bhasa-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages (2023.acl-short)
Copied to clipboard
| Challenge: | Existing tools for language identification are noisy, small and similar to high-resource languages. |
| Approach: | They create a language identification test set for native-script and romanized text which spans all 22 Indic languages and train a model for romanized script. |
| Outcome: | The proposed model improves on native-script and romanized script, and is competitive or better than existing LIDs. |
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages (2022.tacl-1)
Copied to clipboard
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh Shantadevi Khapra
| Challenge: | We present Samanantar, the largest publicly available parallel corpora collection for Indic languages . based on existing corporative, there has been limited benefit for resource-poor languages despite the lack of parallel corporals and monolingual corporata. |
| Approach: | They compile 12.4 million sentence pairs from existing corpora and mine 37.4 million from the Web. |
| Outcome: | The proposed model outperforms existing models and benchmarks on public datasets. |
TRANSLIT: A Large-scale Name Transliteration Resource (2020.lrec-1)
Copied to clipboard
| Challenge: | Transliteration is the process of expressing a proper name from a source language in the characters of a target language. |
| Approach: | They present a large-scale corpus of transliterated names in 180 languages . they use machine learning to train automatic transliteration . |
| Outcome: | The proposed system achieves 92% accuracy on identification of transliterated pairs. |