Papers with Punjabi

14 papers
Towards Speech to Speech Machine Translation focusing on Indian Languages (2023.eacl-demo)

Copied to clipboard

Challenge: SSMT is a web application for translating videos from one language to another by cascading multiple language modules.
Approach: They introduce an SSMT pipeline for translating videos from one language to another by cascading multiple language modules.
Outcome: The proposed system can get 3.5+ MOS score for English to Hindi using human intervention.
Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages (2020.acl-srw)

Copied to clipboard

Challenge: Neural Machine Translation (NMT) is a rapidly advancing MT paradigm that can be used to improve machine translation for many languages.
Approach: They propose a technique called Unified Transliteration and Subword Segmentation to leverage language similarity while exploiting parallel data from related languages.
Outcome: The proposed approach improves translation accuracy by 5 BLEU points over the standard Transformer-based NMT models.
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)

Copied to clipboard

Challenge: IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages.
Approach: They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families.
Outcome: Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages.
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages? (2024.acl-short)

Copied to clipboard

Challenge: a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources.
Approach: They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages.
Outcome: The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi.
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)

Copied to clipboard

Challenge: a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages .
Approach: They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation .
Outcome: The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU .
Por Qué Não Utiliser Alla Språk? Mixed Training with Gradient Optimization in Few-Shot Cross-Lingual Transfer (2022.findings-naacl)

Copied to clipboard

Challenge: a lack of labeled data for low-resource languages leads to the need for effective cross-lingual transfer learning.
Approach: They propose a mixed training method that trains on both source and target data with stochastic gradient surgery, a novel gradient-level optimization.
Outcome: The proposed method outperforms current methods on all tasks and escapes overfitting issues.
Script-Agnosticism and its Impact on Language Identification for Dravidian Languages (2025.naacl-long)

Copied to clipboard

Challenge: a recent study shows that modern systems are script-dependent in language identification (langID) many languages are written in multiple writing systems, and script diversity is common in low-resource languages.
Approach: They propose to learn script-agnostic representations using different strategies . they use word-level script randomization and script exposure to a language written in multiple scripts .
Outcome: The proposed methods exploit script randomization and exposure to a language written in multiple scripts to improve language identification while maintaining competitive performance on naturally occurring text.
Challenge Dataset of Cognates and False Friend Pairs from Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation.
Approach: They create two cognate datasets for twelve Indian languages and use them to generate cognate sets.
Outcome: The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers.
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
Universal Dependencies for Punjabi (2022.lrec-1)

Copied to clipboard

Challenge: UD is a community project that maintains a standard scheme for the annotation of grammar in a cross-lingually consistent manner.
Approach: They propose a Universal Dependencies treebank for Punjabi written in the Gurmukhi script and discuss corpus design and linguistic phenomena encountered in annotation.
Outcome: The proposed treebank covers a variety of genres and has been annotated for POS tags, dependency relations, and graph-based Enhanced Dependencies.
Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages (2021.emnlp-main)

Copied to clipboard

Challenge: A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning.
Approach: They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning.
Outcome: The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning.
Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to predict schwa deletion in Hindi are based on prosodic or phonetic analysis.
Approach: They propose to use Hindi grapheme-to-phoneme (G2P) conversion to predict whether a schwa represented in the orthography is pronounced or unpronounced (deleted).
Outcome: The proposed model outperforms existing models on a newly-compiled pronunciation lexicon extracted from various online dictionaries.
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages.
Approach: They propose to use global benchmarks to track progress on all NLP datasets in all languages.
Outcome: a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench.
Framing Political Bias in Multilingual LLMs Across Pakistani Languages (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) shape public discourse, yet most evaluations of economic and political bias focus on high-resource Western languages and contexts.
Approach: They propose to use a culturally adapted Political Compass Test to evaluate political bias in 13 state-of-the-art LLMs across five Pakistani languages.
Outcome: The proposed framework captures ideological stance (economic/social axes) and stylistic framing (content, tone, emphasis) in 13 state-of-the-art LLMs across five Pakistani languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations