CPLM, a Parallel Corpus for Mexican Languages: Development and Interface (2020.lrec-1)
Copied to clipboard
| Challenge: | Mexico has 68 linguistic groups and 364 varieties, but lack of data on social media and internet is putting them at risk. |
| Approach: | They propose a collaborative corpus for endangered languages in Mexico . they propose linguistic search, digitalization and alignment process for each language . |
| Outcome: | The proposed corpus aligns Spanish with six indigenous languages: Maya, Ch’ol, Mazatec, Mixtec, Otomi, and Nahuatl. |
Similar Papers
Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl (2025.naacl-long)
Copied to clipboard
| Challenge: | a corpus of audio and annotated transcriptions of an endangered Nahuatl is presented . data made available in this corpus are useful for ASR, spelling normalization, and word-level language identification. |
| Approach: | They present a corpus of audio and annotated transcriptions of an endangered Nahuatl in Mexico . the data are useful for ASR, spelling normalization, and word-level language identification . |
| Outcome: | The corpus is made available for use in ASR, spelling normalization, and word-level language identification tasks. |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)
Copied to clipboard
| Challenge: | 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process. |
| Approach: | They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments . |
| Outcome: | The proposed tool can be set up for any parallel corpus and explores its quality and properties. |
Universal Dependencies for Western Sierra Puebla Nahuatl (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of western Sierra Puebla Nahuatl conforms to universal dependency project annotation guidelines . morphological and syntactic phenomena can be analyzed quantitatively with a large enough corpus . |
| Approach: | They present a morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl . it is the first indigenous language of Mexico to be added to the Universal Dependencies project . UD is a widely-used annotation framework whose aim is to provide a consistent schema for morphological and syntactic phenomena for all of the world's languages. |
| Outcome: | The morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl conforms to the universal dependency project annotation guidelines. |
Multilingual Parallel Corpus for Global Communication Plan (L18-1)
Copied to clipboard
| Challenge: | In this paper, we introduce the Global Communication Plan (GCP) Corpus . the corpus is sentence-aligned and covers ten languages, including many Asian languages . |
| Approach: | They introduce the Global Communication Plan (GCP) Corpus, a multilingual parallel corpus . it is sentence-aligned and covers ten languages, including many Asian languages . |
| Outcome: | The proposed corpus is sentence-aligned and covers ten languages, including many Asian languages. |
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)
Copied to clipboard
| Challenge: | Scielo database contains articles from several research domains. |
| Approach: | They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish. |
| Outcome: | The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish. |
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus (2025.findings-emnlp)
Copied to clipboard
| Challenge: | linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources . |
| Approach: | They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks . |
| Outcome: | The proposed corpus spans critical domains such as education, administration, and news. |
Curated Datasets and Neural Models for Machine Translation of Informal Registers between Mayan and Spanish Vernaculars (2024.naacl-long)
Copied to clipboard
| Challenge: | a set of corpora in several Mayan languages spoken in Guatemala and Mexico is published . the languages are considered to be somewhat in decline in terms of resources and global exposure . |
| Approach: | They develop, curate, and publicly release a set of corpora in several Mayan languages spoken in Guatemala and southern Mexico, which they call MayanV. |
| Outcome: | The proposed datasets are parallel with Spanish, the dominant language of the region, and differ in register from most other available resources. |
Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study (2021.acl-long)
Copied to clipboard
Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, Sunita Sarawagi
| Challenge: | Recent research in multilingual language models (LMs) has demonstrated their ability to effectively handle multiple languages in a single model. |
| Approach: | They propose to exploit relatedness among languages in a language family to overcome corpora limitations of LRLs. |
| Outcome: | The proposed model exploits relatedness among languages in a language family to overcome corpora limitations for low web-resource languages. |