Chahta Anumpa: A multimodal corpus of the Choctaw Language (L18-1)

Copied to clipboard

Challenge: a corpus of texts representing the Choctaw language is presented for use in linguistic studies.
Approach: They present a general use corpus for the Choctaw language in the southeastern u.s. the corpus contains audio, video, and text resources, with many texts also translated in english.
Outcome: The proposed corpus provides documentation support for the threatened language . the data set includes audio, video, and text resources .

Similar Papers

Exploring a Choctaw Language Corpus with Word Vectors and Minimum Distance Length (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to explore low resource languages that require no expert knowledge or substantial labor are limited.
Approach: They introduce additions to the Choctaw corpus by using off-the-shelf tools word2vec and Linguistica to create new computational resources for the American indigenous language.
Outcome: The proposed tools can be implemented with minimal labor in the American indigenous language Choctaw.
Cree Corpus: A Collection of nêhiyawêwin Resources (2022.acl-long)

Copied to clipboard

Challenge: Plains Cree is a low resource language with no corpus available for development . a lack of publicly available corpora hinders the development of such technologies .
Approach: They develop a corpus of Plains Cree (nêhiyawêwin) covering genres, time periods, and texts for a variety of intended audiences.
Outcome: The corpus covers genres, time periods, and texts for a variety of intended audiences.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl (2025.naacl-long)

Copied to clipboard

Challenge: a corpus of audio and annotated transcriptions of an endangered Nahuatl is presented . data made available in this corpus are useful for ASR, spelling normalization, and word-level language identification.
Approach: They present a corpus of audio and annotated transcriptions of an endangered Nahuatl in Mexico . the data are useful for ASR, spelling normalization, and word-level language identification .
Outcome: The corpus is made available for use in ASR, spelling normalization, and word-level language identification tasks.
The Abkhaz National Corpus (L18-1)

Copied to clipboard

Challenge: Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language .
Approach: They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus .
Outcome: The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives.
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages.
Approach: They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks.
Outcome: The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language .
Towards Language Technology for Mi’kmaq (L18-1)

Copied to clipboard

Challenge: Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada .
Approach: They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages .
Outcome: The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue .
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)

Copied to clipboard

Challenge: CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting.
Approach: They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages.
Outcome: The proposed model is based on an hour of annotated data and is usable by linguists.
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)

Copied to clipboard

Challenge: Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems.
Approach: They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples.
Outcome: The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations