Challenge: a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi .
Approach: They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines .
Outcome: The proposed corpus provides a language id at two levels: word and sentence.

Similar Papers

MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Sociolinguists and psychologists have been studying these variations in the lexicons and the language from the 50's . code-mixing is a popular method for understanding people's emotions and attitudes towards various subjects, but low-resourced languages often have a mix of scripts and languages.
Approach: They introduce a new sentiment data, MaCMS, for Magahi-Hindi-English code-mixed language, where Magai is a less-resourced minority language.
Outcome: The proposed dataset is the first Magahi-Hindi-English code-mixed dataset for sentiment analysis tasks.
SyMCoM - Syntactic Measure of Code Mixing A Study Of English-Hindi Code-Mixing (2022.findings-acl)

Copied to clipboard

Challenge: Recent work on code mixing in computational settings has leveraged social media code mixed texts to train NLP models.
Approach: They propose to use language ID tags to measure syntactic variety in code-mixed text and their relationship with computational model performance.
Outcome: The proposed measure can be applied to English(en)-hindi(hi) code-mixed datasets and compares them with other measures.
MUTANT: A Multi-sentential Code-mixed Hinglish Dataset (2023.findings-eacl)

Copied to clipboard

Challenge: Existing methods to identify code-mixed text are difficult to scale effectively and efficiently on multi-sentential data.
Approach: They propose to identify multi-sentential code-mixed text (MCT) from multilingual articles using a token-level language-aware pipeline.
Outcome: The proposed dataset includes 67k articles with 85k identified Hinglish MCTs.
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing (2025.findings-emnlp)

Copied to clipboard

Challenge: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances across five core NLP tasks are annotating by three bilingual annotators .
Approach: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances are annotating by three bilingual annotators .
Outcome: The dataset covers five core NLP tasks, including Token-level Language Identification, Matrix Language Identification and Named Entity Recognition.
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien (2022.findings-emnlp)

Copied to clipboard

Challenge: CM is a challenging task when mixed languages include dialects.
Approach: They propose to construct a Hokkien-Mandarin CM dataset to overcome the limitation . they propose to use a linguistics-based toolkit to train the model for translation tasks .
Outcome: The proposed model achieves good results on CM data translation while maintaining monolingual translation quality.
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)

Copied to clipboard

Challenge: a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet .
Approach: They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages.
Outcome: The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset .
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)

Copied to clipboard

Challenge: Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation.
Approach: They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi.
Outcome: The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge.
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing web-mined datasets for low-resource languages have been useful for low resource NLP.
Approach: They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable.
Outcome: The proposed model outperforms baselines when balancing F1 and false positive rate (FPR).
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus (2025.findings-emnlp)

Copied to clipboard

Challenge: linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources .
Approach: They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks .
Outcome: The proposed corpus spans critical domains such as education, administration, and news.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations