Challenge: Mexico has 68 linguistic groups and 364 varieties, but lack of data on social media and internet is putting them at risk.
Approach: They propose a collaborative corpus for endangered languages in Mexico . they propose linguistic search, digitalization and alignment process for each language .
Outcome: The proposed corpus aligns Spanish with six indigenous languages: Maya, Ch’ol, Mazatec, Mixtec, Otomi, and Nahuatl.

Similar Papers

Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl (2025.naacl-long)

Copied to clipboard

Challenge: a corpus of audio and annotated transcriptions of an endangered Nahuatl is presented . data made available in this corpus are useful for ASR, spelling normalization, and word-level language identification.
Approach: They present a corpus of audio and annotated transcriptions of an endangered Nahuatl in Mexico . the data are useful for ASR, spelling normalization, and word-level language identification .
Outcome: The corpus is made available for use in ASR, spelling normalization, and word-level language identification tasks.
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)

Copied to clipboard

Challenge: 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process.
Approach: They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments .
Outcome: The proposed tool can be set up for any parallel corpus and explores its quality and properties.
Universal Dependencies for Western Sierra Puebla Nahuatl (2022.lrec-1)

Copied to clipboard

Challenge: Annotated corpus of western Sierra Puebla Nahuatl conforms to universal dependency project annotation guidelines . morphological and syntactic phenomena can be analyzed quantitatively with a large enough corpus .
Approach: They present a morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl . it is the first indigenous language of Mexico to be added to the Universal Dependencies project . UD is a widely-used annotation framework whose aim is to provide a consistent schema for morphological and syntactic phenomena for all of the world's languages.
Outcome: The morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl conforms to the universal dependency project annotation guidelines.
Multilingual Parallel Corpus for Global Communication Plan (L18-1)

Copied to clipboard

Challenge: In this paper, we introduce the Global Communication Plan (GCP) Corpus . the corpus is sentence-aligned and covers ten languages, including many Asian languages .
Approach: They introduce the Global Communication Plan (GCP) Corpus, a multilingual parallel corpus . it is sentence-aligned and covers ten languages, including many Asian languages .
Outcome: The proposed corpus is sentence-aligned and covers ten languages, including many Asian languages.
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)

Copied to clipboard

Challenge: Scielo database contains articles from several research domains.
Approach: They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish.
Outcome: The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish.
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus (2025.findings-emnlp)

Copied to clipboard

Challenge: linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources .
Approach: They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks .
Outcome: The proposed corpus spans critical domains such as education, administration, and news.
Curated Datasets and Neural Models for Machine Translation of Informal Registers between Mayan and Spanish Vernaculars (2024.naacl-long)

Copied to clipboard

Challenge: a set of corpora in several Mayan languages spoken in Guatemala and Mexico is published . the languages are considered to be somewhat in decline in terms of resources and global exposure .
Approach: They develop, curate, and publicly release a set of corpora in several Mayan languages spoken in Guatemala and southern Mexico, which they call MayanV.
Outcome: The proposed datasets are parallel with Spanish, the dominant language of the region, and differ in register from most other available resources.
Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study (2021.acl-long)

Copied to clipboard

Challenge: Recent research in multilingual language models (LMs) has demonstrated their ability to effectively handle multiple languages in a single model.
Approach: They propose to exploit relatedness among languages in a language family to overcome corpora limitations of LRLs.
Outcome: The proposed model exploits relatedness among languages in a language family to overcome corpora limitations for low web-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations