Papers with Telugu
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication (2025.coling-main)
Copied to clipboard
| Challenge: | Existing research indicates that disfluencies can constitute up to 5.9% of words in spontaneous speech, with repetitions accounting for over half of these disfluency. |
| Approach: | They propose to use a dataset to analyze reduplication and repetition in speech using computational linguistics to evaluate transformer-based models. |
| Outcome: | The proposed models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification. |
Towards Speech to Speech Machine Translation focusing on Indian Languages (2023.eacl-demo)
Copied to clipboard
| Challenge: | SSMT is a web application for translating videos from one language to another by cascading multiple language modules. |
| Approach: | They introduce an SSMT pipeline for translating videos from one language to another by cascading multiple language modules. |
| Outcome: | The proposed system can get 3.5+ MOS score for English to Hindi using human intervention. |
A Simple and Effective Dependency Parser for Telugu (2020.acl-srw)
Copied to clipboard
| Challenge: | Existing dependency parsers for Telugu use hand-crafted features based on linguistic information like part-of-speech and morphology which are expensive to annotate. |
| Approach: | They propose to replace linguistic feature templates with a minimal feature function for Telugu . they train a BERT model on the Telugus Wikipedia data and use contextual vector representations to train the parser. |
| Outcome: | The proposed parser achieves state-of-the-art for Telugu using contextual vector representations . the proposed model trains on the Telugus Wikipedia data and trains with a greedy transition based approach . |
Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages (2020.acl-srw)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a rapidly advancing MT paradigm that can be used to improve machine translation for many languages. |
| Approach: | They propose a technique called Unified Transliteration and Subword Segmentation to leverage language similarity while exploiting parallel data from related languages. |
| Outcome: | The proposed approach improves translation accuracy by 5 BLEU points over the standard Transformer-based NMT models. |
Samvaadhana: A Telugu Dialogue System in Hospital Domain (D19-61)
Copied to clipboard
| Challenge: | a dialogue system for Hospital domain in Telugu is a resource-poor Dravidian language . the system handles various hospital and doctor related queries . |
| Approach: | They propose to model a dialogue system for Hospital domain in Telugu which is a resource-poor Dravidian language. |
| Outcome: | The proposed system achieves a high overall rating and a significantly accurate context-capturing method. |
Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios (2026.eacl-short)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) systems for low-resource languages produce erroneous transcripts due to limited annotated data and linguistic complexity. |
| Approach: | They compare language models and large language models for post-ASR correction in Hindi . they observe a scaling trend under zero-shot ICL where mid-sized LLMs degrade performance before marginal recovery at extreme scales. |
| Outcome: | The proposed model outperforms larger models in both fine-tuning and in-context learning settings. |
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing vision language navigation tasks require soft attention over words to locate instructions . a new approach uses syntax information to ground instructions with visual information . |
| Approach: | They propose a vision language navigation agent that utilizes syntax information to enhance alignment between the instruction and the current visual scenes. |
| Outcome: | The proposed agent outperforms the baseline model that does not use syntax information on the Room-to-Room dataset, especially in the unseen environment. |
Resource Creation Towards Automated Sentiment Analysis in Telugu (a low resource language) and Integrating Multiple Domain Sources to Enhance Sentiment Prediction (L18-1)
Copied to clipboard
| Challenge: | Sentiment Analysis of text is an important task in many applications . but the task becomes challenging when it comes to low resource languages . |
| Approach: | They propose to create a corpus of polarity-based sentiment classifiers in Telugu for different domains like movie reviews, song lyrics, product reviews and book reviews. |
| Outcome: | The proposed model performs well in multiple domains and is compared with the previous models. |
Bidirectional Reasoning Supervision for Multilingual Financial Decision Making (2025.emnlp-industry)
Copied to clipboard
Muhammad Rafsan Kabir, Jawad Ibn Ahad, Robin Krambroeckers, Silvia Ahmed, M M Lutfe Elahi, Nabeel Mohammed, Shafin Rahman
| Challenge: | Large Language Models have been used for sentiment analysis, machine translation, and question answering, but their effectiveness in the multilingual financial domain remains unknown. |
| Approach: | They propose a fine-tuning approach that integrates positive and negative rationales alongside classification labels. |
| Outcome: | The proposed approach outperforms existing methods across English, Hindi, Bengali, and Telugu, and is suitable for industry applications. |
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)
Copied to clipboard
Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
Intent Identification and Entity Extraction for Healthcare Queries in Indic Languages (2023.findings-eacl)
Copied to clipboard
| Challenge: | Currently, there is a lack of data and technology for resource-poor languages in developing countries like India. |
| Approach: | They propose to use two different datasets to analyze query intents and entities in healthcare. |
| Outcome: | The proposed model is useful to identify query intents and entities in real-world scenarios. |
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)
Copied to clipboard
Nedjma Ousidhoum, Shamsuddeen Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Ahmad, Sanchit Ahuja, Alham Aji, Vladimir Araujo, Abinew Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine Kock, Genet Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Tilaye, Krishnapriya Vishnubhotla, Genta Winata, Seid Yimam, Saif Mohammad
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding (2020.emnlp-main)
Copied to clipboard
| Challenge: | Room-Across-Room (RxR) is a vision-and-language navigation dataset that addresses gaps in existing ones by addressing known biases in paths and eliciting more references to visible entities. |
| Approach: | They introduce a new Vision-and-Language Navigation (VLN) dataset that addresses biases in paths and elicits more references to visible entities. |
| Outcome: | The proposed model learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. |
Challenge Dataset of Cognates and False Friend Pairs from Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation. |
| Approach: | They create two cognate datasets for twelve Indian languages and use them to generate cognate sets. |
| Outcome: | The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers. |
Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati (2024.lrec-main)
Copied to clipboard
| Challenge: | There has been a rise in homophobic and transphobic content targeting LGBT+ individuals on social media platforms. |
| Approach: | They propose to use a dataset to automatically identify homophobic and transphobic content within comments collected from YouTube for three languages. |
| Outcome: | The proposed dataset will identify homophobic and transphobic content within comments collected from YouTube in Telugu, Kannada, and Gujarati. |
HIT - A Hierarchically Fused Deep Attention Network for Robust Code-mixed Language Representation (2021.findings-acl)
Copied to clipboard
| Challenge: | linguistics and morphology of resource-short code-mixed texts remain a key challenge in text processing. |
| Approach: | They propose a hierarchical transformer-based framework that captures the semantic relationship among words and hierarchically learns sentencelevel semantics using a fused attention mechanism. |
| Outcome: | The proposed framework improves on one European and five Indic languages on four NLP tasks on eleven datasets. |
Automatic Speech Recognition in Sanskrit: A New Speech Corpus and Modelling Insights (2021.findings-acl)
Copied to clipboard
| Challenge: | In this paper, we propose the first large scale study of automatic speech recognition in Sanskrit . we focus on the impact of unit selection in San's ASR systems . |
| Approach: | They propose a large scale study of automatic speech recognition in Sanskrit . they propose syllable level unit selection that captures character sequences . |
| Outcome: | The proposed model captures character sequences from one vowel in the word to the next vowela. |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
Dataset Creation and Evaluation of Aspect Based Sentiment Analysis in Telugu, a Low Resource Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Aspect Based Sentiment Analysis (ABSA) is a finer level sentiment analysis that assigns polarity to each targeted aspect instead of the entire review. |
| Approach: | They propose to use Telugu as a language for aspect based sentiment analysis . they use a resource that can be used to classify and categorise aspects of a review . |
| Outcome: | The proposed resource is based on a set of tasks in Telugu which demonstrate its reliability and usefulness. |
IndicFinNLP: Financial Natural Language Processing for Indian Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | IndicFinNLP is a collection of 9 datasets relating to FinNLP for three Indian languages. |
| Approach: | They propose to use financial NLP to detect exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu. |
| Outcome: | The proposed framework detects exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu. |
TeClass: A Human-Annotated Relevance-based Headline Classification and Generation Dataset for Telugu (2024.lrec-main)
Copied to clipboard
| Challenge: | Relevance-based headline classification is under-explored in low-resource languages like Telugu due to a lack of annotated data. |
| Approach: | They propose that relevance-based headline classification can greatly aid the task of generating relevant headlines. |
| Outcome: | The proposed model can generate relevant headlines with 78,534 annotations in Telugu . the model shows a 5 point increment in the ROUGE-L scores . |
Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy (2026.findings-acl)
Copied to clipboard
Vallabhaneni Raj Kumar, Ashwin S, Supriya Manna, Niladri Sett, Cheedella V S N M S Hema Harshitha, Kurakula Harshitha, Basina Deepakraj, Anand Kumar Sharma, Tanuj Sarkar, Samanthapudi Shakeer, Bondada Navaneeth Krishna
| Challenge: | a limited amount of annotated data has slowed progress in machine learning for low-resource languages . a sentiment label records an annotator's final decision, but it is not a valid record of the annotation's interpretation. |
| Approach: | They propose a large-scale Telugu sentiment classification dataset annotated with sentiment labels and human-selected rationales from multiple native speakers. |
| Outcome: | The proposed model improves classification performance, explanation quality, and social bias by incorporating human rationales. |
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages (2026.acl-long)
Copied to clipboard
| Challenge: | Multilingual large language models are expensive to pretrain and suffer from imbalances across languages and datasets. |
| Approach: | They propose a family of Indian language-only autoregressive language models trained on open-source language-specific data for the five most spoken Indian languages. |
| Outcome: | The proposed model outperforms most larger models up to 8B across all five languages. |