Papers with Punjabi
Towards Speech to Speech Machine Translation focusing on Indian Languages (2023.eacl-demo)
Copied to clipboard
| Challenge: | SSMT is a web application for translating videos from one language to another by cascading multiple language modules. |
| Approach: | They introduce an SSMT pipeline for translating videos from one language to another by cascading multiple language modules. |
| Outcome: | The proposed system can get 3.5+ MOS score for English to Hindi using human intervention. |
Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages (2020.acl-srw)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a rapidly advancing MT paradigm that can be used to improve machine translation for many languages. |
| Approach: | They propose a technique called Unified Transliteration and Subword Segmentation to leverage language similarity while exploiting parallel data from related languages. |
| Outcome: | The proposed approach improves translation accuracy by 5 BLEU points over the standard Transformer-based NMT models. |
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages? (2024.acl-short)
Copied to clipboard
| Challenge: | a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources. |
| Approach: | They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages. |
| Outcome: | The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi. |
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)
Copied to clipboard
Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
Por Qué Não Utiliser Alla Språk? Mixed Training with Gradient Optimization in Few-Shot Cross-Lingual Transfer (2022.findings-naacl)
Copied to clipboard
| Challenge: | a lack of labeled data for low-resource languages leads to the need for effective cross-lingual transfer learning. |
| Approach: | They propose a mixed training method that trains on both source and target data with stochastic gradient surgery, a novel gradient-level optimization. |
| Outcome: | The proposed method outperforms current methods on all tasks and escapes overfitting issues. |
Script-Agnosticism and its Impact on Language Identification for Dravidian Languages (2025.naacl-long)
Copied to clipboard
| Challenge: | a recent study shows that modern systems are script-dependent in language identification (langID) many languages are written in multiple writing systems, and script diversity is common in low-resource languages. |
| Approach: | They propose to learn script-agnostic representations using different strategies . they use word-level script randomization and script exposure to a language written in multiple scripts . |
| Outcome: | The proposed methods exploit script randomization and exposure to a language written in multiple scripts to improve language identification while maintaining competitive performance on naturally occurring text. |
Challenge Dataset of Cognates and False Friend Pairs from Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation. |
| Approach: | They create two cognate datasets for twelve Indian languages and use them to generate cognate sets. |
| Outcome: | The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers. |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
Universal Dependencies for Punjabi (2022.lrec-1)
Copied to clipboard
| Challenge: | UD is a community project that maintains a standard scheme for the annotation of grammar in a cross-lingually consistent manner. |
| Approach: | They propose a Universal Dependencies treebank for Punjabi written in the Gurmukhi script and discuss corpus design and linguistic phenomena encountered in annotation. |
| Outcome: | The proposed treebank covers a variety of genres and has been annotated for POS tags, dependency relations, and graph-based Enhanced Dependencies. |
Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning. |
| Approach: | They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning. |
| Outcome: | The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning. |
Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods to predict schwa deletion in Hindi are based on prosodic or phonetic analysis. |
| Approach: | They propose to use Hindi grapheme-to-phoneme (G2P) conversion to predict whether a schwa represented in the orthography is pronounced or unpronounced (deleted). |
| Outcome: | The proposed model outperforms existing models on a newly-compiled pronunciation lexicon extracted from various online dictionaries. |
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)
Copied to clipboard
Yueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig
| Challenge: | despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages. |
| Approach: | They propose to use global benchmarks to track progress on all NLP datasets in all languages. |
| Outcome: | a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench. |
Framing Political Bias in Multilingual LLMs Across Pakistani Languages (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) shape public discourse, yet most evaluations of economic and political bias focus on high-resource Western languages and contexts. |
| Approach: | They propose to use a culturally adapted Political Compass Test to evaluate political bias in 13 state-of-the-art LLMs across five Pakistani languages. |
| Outcome: | The proposed framework captures ideological stance (economic/social axes) and stylistic framing (content, tone, emphasis) in 13 state-of-the-art LLMs across five Pakistani languages. |