| Challenge: | a corpus of texts representing the Choctaw language is presented for use in linguistic studies. |
| Approach: | They present a general use corpus for the Choctaw language in the southeastern u.s. the corpus contains audio, video, and text resources, with many texts also translated in english. |
| Outcome: | The proposed corpus provides documentation support for the threatened language . the data set includes audio, video, and text resources . |
Similar Papers
Exploring a Choctaw Language Corpus with Word Vectors and Minimum Distance Length (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools to explore low resource languages that require no expert knowledge or substantial labor are limited. |
| Approach: | They introduce additions to the Choctaw corpus by using off-the-shelf tools word2vec and Linguistica to create new computational resources for the American indigenous language. |
| Outcome: | The proposed tools can be implemented with minimal labor in the American indigenous language Choctaw. |
Cree Corpus: A Collection of nêhiyawêwin Resources (2022.acl-long)
Copied to clipboard
| Challenge: | Plains Cree is a low resource language with no corpus available for development . a lack of publicly available corpora hinders the development of such technologies . |
| Approach: | They develop a corpus of Plains Cree (nêhiyawêwin) covering genres, time periods, and texts for a variety of intended audiences. |
| Outcome: | The corpus covers genres, time periods, and texts for a variety of intended audiences. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl (2025.naacl-long)
Copied to clipboard
| Challenge: | a corpus of audio and annotated transcriptions of an endangered Nahuatl is presented . data made available in this corpus are useful for ASR, spelling normalization, and word-level language identification. |
| Approach: | They present a corpus of audio and annotated transcriptions of an endangered Nahuatl in Mexico . the data are useful for ASR, spelling normalization, and word-level language identification . |
| Outcome: | The corpus is made available for use in ASR, spelling normalization, and word-level language identification tasks. |
The Abkhaz National Corpus (L18-1)
Copied to clipboard
| Challenge: | Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language . |
| Approach: | They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus . |
| Outcome: | The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives. |
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)
Copied to clipboard
| Challenge: | Using the corpus, we study the characteristics of interpreters' work and train machine translation systems. |
| Approach: | They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work. |
| Outcome: | The proposed corpus can be used for teaching interpreters and to train machine translation systems. |
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages. |
| Approach: | They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks. |
| Outcome: | The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language . |
Towards Language Technology for Mi’kmaq (L18-1)
Copied to clipboard
| Challenge: | Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada . |
| Approach: | They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages . |
| Outcome: | The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue . |
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)
Copied to clipboard
| Challenge: | CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting. |
| Approach: | They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages. |
| Outcome: | The proposed model is based on an hour of annotated data and is usable by linguists. |
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |