| Challenge: | a new computational task supports the construction of high quality texts and lexicons for low resource languages. |
| Approach: | They propose a computational task which is tuned to the available knowledge and interests in an Indigenous community. |
| Outcome: | The proposed method achieves a transcription density gain of 17% in a morphologically complex language . the proposed grammar includes a description of the phonology and morphosyntax . |
Similar Papers
Enabling Interactive Transcription in an Indigenous Community (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for manual transcription are often in isolation from the speech community, and so we miss out on the opportunity to take advantage of the interests and skills of local people. |
| Approach: | They propose a transcription workflow which combines spoken term detection and human-in-the-loop to support speech transcription in almost-zero resource settings. |
| Outcome: | The proposed workflow is based on two endangered languages with zero-resource datasets. |
Interactive Word Completion for Morphologically Complex Languages (2020.coling-main)
Copied to clipboard
| Challenge: | morphologically complex languages have multiple morph slots with large or unbounded sets of fillers. |
| Approach: | They propose a method for morphologically-aware text input in Kunwinjku . they modify an existing finite state recognizer to map input morph prefixes to morph completions . |
| Outcome: | The proposed method is portable to Turkish and shows that it can be used in other languages. |
Learning From Failure: Data Capture in an Australian Aboriginal Community (2022.acl-long)
Copied to clipboard
| Challenge: | a prototype of a language data capture app for speakers was tested in an Aboriginal community . elicitation of word lists, phrases, etc. has been used for decades to collect data for Indigenous languages . many software tools are developed to support linguists' work . |
| Approach: | They propose to deploy an app for speakers to confirm system guesses in an approach to transcription based on word spotting. |
| Outcome: | The proposed app was tested in an Aboriginal community in australia . it was able to confirm system guesses without a transcription bottleneck . the results were compared with other apps in the community . |
Morphological Processing of Low-Resource Languages: Where We Are and What’s Next (2022.findings-acl)
Copied to clipboard
Adam Wiemerslage, Miikka Silfverberg, Changbing Yang, Arya McCarthy, Garrett Nicolai, Eliana Colunga, Katharina Kann
| Challenge: | Existing models for morphological processing are not suitable for low-resource languages, but they are still lacking in the field of computational morphology. |
| Approach: | They propose to bridge two unsupervised models to understand a language’s morphology from raw text alone and propose to use them to improve their models. |
| Outcome: | The proposed models perform reasonably, but there is room for improvement. |
KinyaBERT: a Morphology-aware Kinyarwanda Language Model (2022.acl-long)
Copied to clipboard
| Challenge: | Pre-trained language models such as BERT are sub-optimal at handling morphologically rich languages. |
| Approach: | They propose a two-tier BERT architecture that leverages a morphological analyzer and explicitly represents morphology in a low-resource Kinyarwanda language. |
| Outcome: | The proposed model outperforms baseline models on the low-resource morphologically rich Kinyarwanda language by 2% in F1 score and 4.3% in average score of GLUE benchmark. |
Interactive Word Completion for Plains Cree (2022.acl-long)
Copied to clipboard
| Challenge: | a tool that helps users incrementally build complex words is being developed in morphologically complex languages. |
| Approach: | They propose a finite state approach which maps prefixes in a language to completions up to the next morpheme boundary for incremental building of complex words. |
| Outcome: | The proposed approach shows portability to a larger, more complete morphological transducer. |
Empowering Low-Resource Regional Languages with Lexicons : A Comparative Study of NLP Tools for Morphosyntactic Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | a lack of human and financial resources makes integrating lexicon information to low-resource languages challenging. |
| Approach: | They propose to use a bilingual lexicon to integrate lexical information to low-resource language . they compare a lexiconal approach to a neural approach that uses a larger lexicone . |
| Outcome: | The proposed approach improves POS tagging while using different lexicon sizes. |
Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community (2022.coling-1)
Copied to clipboard
| Challenge: | Recent research has focused on low-resource languages and the transcription bottleneck paradigm. |
| Approach: | They propose to use a spoken term detection system to train a speech recognition system in an Aboriginal community to reach better comprehension and engagement from Aboriginal participants. |
| Outcome: | The proposed system can be implemented in an Aboriginal community and reach better comprehension and engagement from Aboriginal participants. |
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)
Copied to clipboard
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noel Kouarata, Lori Lamel, Hélène Maynard, Markus Mueller, Annie Rialland, Sebastian Stueker, François Yvon, Marcely Zanon-Boito
| Challenge: | a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels . |
| Approach: | They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations . |
| Outcome: | The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task . |
Automatic Transcription Challenges for Inuktitut, a Low-Resource Polysynthetic Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Inuktitut is one of the 60 Indigenous languages currently spoken in Canada . polysynthetic languages are often termed agglutinative when their morphemes have clear boundaries and thus are easily segmentable. |
| Approach: | They propose to use a corpus of 23 hours of transcribed oral stories to train automatic speech recognition in Inuktitut. |
| Outcome: | The proposed model shows that Inuktitut displays a much higher degree of polysynthesis than other agglutinative languages like Finnish or Turkish. |