Papers by Anoop Kunchukuttan
IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages (2022.emnlp-main)
Copied to clipboard
Aman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra, Pratyush Kumar
| Challenge: | IndicNLG is a non-English language that is hampered by the scarcity of datasets. |
| Approach: | They propose to create a dataset for natural language generation for 11 Indic languages . they use a set of pre-trained models to train multilingual models . |
| Outcome: | The proposed datasets show that pre-trained models perform well in multilingual and monolingual tasks. |
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)
Copied to clipboard
| Challenge: | We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script. |
| Approach: | They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic. |
| Outcome: | The proposed training recipe improves multilingual transliteration for Indic languages. |
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation (2024.lrec-main)
Copied to clipboard
| Challenge: | a number of languages are used in online conversations, resulting in code-mixing . the problem is largely unexplored due to the lack of annotated data and noise . |
| Approach: | They propose a robust perturbation-based joint-training model that learns to handle noise in code-mixed text by parameter sharing across clean and noisy words. |
| Outcome: | The proposed model learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words. |
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages? (2024.acl-short)
Copied to clipboard
| Challenge: | a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources. |
| Approach: | They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages. |
| Outcome: | The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi. |
Multilingual Neural Machine Translation (2020.coling-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will cover the latest advances in NMT to enhance low-resource translation. |
| Approach: | They will cover the latest advances in NMT approaches that leverage multilingualism . they will focus on topics such as language divergence, transfer learning and pivoting . |
| Outcome: | This tutorial will cover the latest advances in NMT to enhance low-resource translation models. |
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI (2026.eacl-short)
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) are highly effective on mathematical, scientific, and other question-answering tasks. |
| Approach: | They compare an LRM's reasoning in English to that of a multilingual question . they find that English reasoning traces exhibit a substantially higher presence of cognitive behaviors . |
| Outcome: | The LRMs generate reasoning sequences in English, but the language of the question is not. |
Naamapadam: A Large-Scale Named Entity Annotated Data for Indic Languages (2023.acl-long)
Copied to clipboard
Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, Pratyush Kumar, Rudra Murthy, Anoop Kunchukuttan
| Challenge: | Named Entity Recognition (NER) is a fundamental task in natural language processing (NLP). |
| Approach: | They present the largest publicly available Named Entity Recognition dataset for the 11 major Indian languages from two language families. |
| Outcome: | The proposed dataset is the largest publicly available Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families. |
Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users (2023.findings-emnlp)
Copied to clipboard
Yash Madhani, Sushane Parthan, Priyanka Bedekar, Gokul Nc, Ruchi Khapra, Anoop Kunchukuttan, Pratyush Kumar, Mitesh Khapra
| Challenge: | Indian subcontinent is home to diverse languages written in multiple scripts . widespread use of romanization and lack of standardization means accurate transliteration models form a critical component in the NLP stack for Indian languages used by over 735 million Internet users. |
| Approach: | They propose to build a transliteration dataset using monolingual and parallel corpora and human annotators. |
| Outcome: | The proposed model improves accuracy by 15% on the Dakshina test set and establishes strong baselines on the Aksharantar test set. |
Data and Model Centric Approaches for Expansion of Large Language Models to New languages (2025.emnlp-tutorials)
Copied to clipboard
| Challenge: | Existing LLMs mainly support English alongside a handful of high resource languages . this leaves a major gap for most low-resource languages despite increasing pace of research . |
| Approach: | This tutorial examines approaches to expand the language coverage of LLMs . they look at tokenizer training, pre-training, instruction tuning, alignment, evaluation, etc. |
| Outcome: | This tutorial examines approaches to expand the language coverage of LLMs . it provides an efficient and viable path to bring LLM technologies to low-resource languages . |
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages (2020.findings-emnlp)
Copied to clipboard
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, Pratyush Kumar
| Challenge: | In this paper, we present NLP resources for 11 major Indian languages . distributional representations are the cornerstone of modern NLP, authors say . |
| Approach: | They introduce NLP resources for 11 major Indian languages from two major language families . monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . they also compile a benchmark for Indian language NLU to evaluate their results . |
| Outcome: | The monolingual corpora contains 8.8 billion tokens across all 11 languages and Indian English . the pre-trained language models are based on the compact ALBERT model . |
Addressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource Languages (N19-1)
Copied to clipboard
| Challenge: | Existing studies show that transfer learning works best when the languages are related. |
| Approach: | They propose to pre-order assisting language sentences to match the word order of the source language and train the parent model. |
| Outcome: | The proposed model can improve translation quality in low-resource scenarios by pre-ordering the assisting language sentences to match the word order of the source language and training the parent model. |
Bilingual Tabular Inference: A Case Study on Indic Languages (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing studies on Tabular Natural Language Inference (TNLI) focus on monolingual settings where tabular premise and hypothesis are in the same language. |
| Approach: | They propose a task where tabular premise and hypothesis are in two languages . they translate textual hypotheses from an English-indic TNLI dataset into eleven major languages - english and indic . |
| Outcome: | The proposed model performs well on a bilingual dataset in English and in 11 major Indian languages. |
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages (2022.tacl-1)
Copied to clipboard
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh Shantadevi Khapra
| Challenge: | We present Samanantar, the largest publicly available parallel corpora collection for Indic languages . based on existing corporative, there has been limited benefit for resource-poor languages despite the lack of parallel corporals and monolingual corporata. |
| Approach: | They compile 12.4 million sentence pairs from existing corpora and mine 37.4 million from the Web. |
| Outcome: | The proposed model outperforms existing models and benchmarks on public datasets. |
Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages (2023.acl-long)
Copied to clipboard
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, Pratyush Kumar
| Challenge: | Recent advances in Natural Language Understanding are driven by pretrained multilingual models, which can potentially reduce the performance gap between high-resource languages through zero-shot knowledge transfer. |
| Approach: | They propose to create a human-supervised benchmark for Indic languages, IndicXTREME, with nine diverse NLU tasks covering 20 languages. |
| Outcome: | The proposed model improves on the monolingual corpora, IndicCorp, and IndicBERT in Indic languages with 105 evaluation sets across languages and tasks. |
IndicXNLI: Evaluating Multilingual Inference for Indian Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | Indic NLP has made rapid advances in terms of corpora and pre-trained models, but benchmark datasets on standard NLU tasks are limited. |
| Approach: | They propose to use an NLI dataset for 11 Indic languages to test their accuracy. |
| Outcome: | The proposed dataset provides useful insights into the behaviour of pre-trained models for a diverse set of languages. |
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)
Copied to clipboard
Ashwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M Khapra, Raj Dabre
| Challenge: | Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks. |
| Approach: | They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better . |
| Outcome: | The proposed model improves on existing models and is open-source with permissive licenses. |
The IIT Bombay English-Hindi Parallel Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of 1.49 million parallel segments is available in the public domain . the corpus is the largest publicly available English-Hindi parallel corpus . |
| Approach: | They present the IIT Bombay English-Hindi Parallel Corpus . they present a compilation of public and private parallel corpora . |
| Outcome: | The corpus contains 1.49 million parallel segments, of which 694k were not previously available in the public domain. |
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit strong multilingual performance despite training on English-centric corpora. |
| Approach: | They propose to use Romanization as a potential bridge in multilingual processing . they propose to encode semantic concepts similarly across native and Romanized scripts . |
| Outcome: | The proposed model encodes semantic concepts across native and Romanized scripts, suggesting a shared underlying representation. |
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation. |
| Approach: | They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching. |
| Outcome: | The proposed model performs poorly when faced with reordered constraints or irrelevant information. |
CTQScorer: Combining Multiple Features for In-context Example Selection for Machine Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have demonstrated the capability to perform on machine translation when the input is prompted with a few examples. |
| Approach: | They propose a regression model that combine features influencing example selection to maximize translation quality. |
| Outcome: | The proposed model outperforms random selection and strong single-factor baselines on multiple language pairs and language models. |
Overview of the 6th Workshop on Asian Translation (D19-52)
Copied to clipboard
Toshiaki Nakazawa, Nobushige Doi, Shohei Higashiyama, Chenchen Ding, Raj Dabre, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Yusuke Oda, Shantipriya Parida, Ondřej Bojar, Sadao Kurohashi
| Challenge: | The 6th workshop on Asian translation (WAT2019) was held in hong kong, hongkong, and hong kong. |
| Approach: | They present the results of the shared tasks from the 6th workshop on Asian translation (WAT2019) 25 teams participated in the shared task and 10 research paper submissions were accepted . |
| Outcome: | The results of the 6th workshop on Asian translation (WAT2019) include JaEn, JaZh scientific paper translation subtasks, Ja'En, ja'Ko, Ja’En patent translation sub tasks, Hi'En and My'En patent subtask and Ru'Ja news commentary translation task. |
CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages (2024.eacl-short)
Copied to clipboard
| Challenge: | Existing models for ELRLs lack parallel corpora and monolingual corporata . authors propose novel character-span noise argumentation model to facilitate cross-lingual transfer . |
| Approach: | They propose a character-span noise argumentation model to facilitate cross-lingual transfer . they use character-size noise argumentations to regularize training data of HRL . |
| Outcome: | The proposed model outperforms baselines on closely related HRL-ELRL pairs from three different language families. |
IndicBART: A Pre-trained Model for Indic Natural Language Generation (2022.findings-acl)
Copied to clipboard
| Challenge: | IndicBART is a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages and English. |
| Approach: | They present a multilingual sequence-to-sequence pre-trained model for Indic languages . they evaluate it on two NLG tasks: Neural Machine Translation and extreme summarization . |
| Outcome: | The proposed model performs well on low-resource translation scenarios . Script sharing, multilingual training, and better utilization contribute to the performance. |
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER (P18-2)
Copied to clipboard
| Challenge: | Existing approaches to improve NER performance add training data from one or more assisting languages to the primary language. |
| Approach: | They propose a metric based on symmetric KL divergence to filter out highly divergent training instances in the assisting language. |
| Outcome: | The proposed method improves NER performance in many languages, including those with limited training data. |
IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian Languages (2023.acl-long)
Copied to clipboard
Ananya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre
| Challenge: | Recent studies on machine translation systems focus on high-resource languages, but focus has shifted to low-resourced languages. |
| Approach: | They evaluate 16 metrics from a multidimensional quality metric dataset . they show pre-trained metrics have higher correlations with annotator scores . |
| Outcome: | The proposed evaluations show that pre-trained metrics outperform COMET on Indian languages. |
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs (2025.acl-long)
Copied to clipboard
Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M Khapra
| Challenge: | Evaluating machine-generated text remains a challenge in NLP for non-English languages . current evaluation frameworks focus on English, revealing a gap in multilingual evaluations . |
| Approach: | They propose a cross-lingual auto evaluation framework that includes evaluator LLMs and a test set specifically designed for multilingual evaluation. |
| Outcome: | The proposed model aligns more closely with human judgments than proprietary models on non-English language evaluations. |
Bhasa-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages (2023.acl-short)
Copied to clipboard
| Challenge: | Existing tools for language identification are noisy, small and similar to high-resource languages. |
| Approach: | They create a language identification test set for native-script and romanized text which spans all 22 Indic languages and train a model for romanized script. |
| Outcome: | The proposed model improves on native-script and romanized script, and is competitive or better than existing LIDs. |
DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent work shows the power of few-shot prompting with large language models for tasks like machine translation, summarization, and question answering. |
| Approach: | They propose a few-shot prompting approach that decomposes the translation process into word chunks. |
| Outcome: | The proposed approach outperforms established few-shot prompting models with 8 chrF++ scores across languages. |