Towards Language Technology for Mi’kmaq (L18-1)

Copied to clipboard

Challenge: Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada .
Approach: They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages .
Outcome: The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue .

Similar Papers

Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi’kmaq Language Modelling (2020.lrec-1)

Copied to clipboard

Challenge: Mi'kmaq is an Indigenous language spoken primarily in Eastern Canada.
Approach: They consider n-gram and RNN language models for Mi'kmaq and use them to investigate their performance.
Outcome: The proposed model performs better than word-level models, but does not improve over word-based models.
Indigenous language technologies in Canada: Assessment, challenges, and successes (C18-1)

Copied to clipboard

Challenge: There are approximately 60 Indigenous languages currently spoken in Canada.
Approach: They examine which technologies have been developed and which are feasible to develop for the 60 Indigenous languages spoken in Canada.
Outcome: The proposed technologies are based on the existing technologies and are feasible for most or all of these languages.
Cree Corpus: A Collection of nêhiyawêwin Resources (2022.acl-long)

Copied to clipboard

Challenge: Plains Cree is a low resource language with no corpus available for development . a lack of publicly available corpora hinders the development of such technologies .
Approach: They develop a corpus of Plains Cree (nêhiyawêwin) covering genres, time periods, and texts for a variety of intended audiences.
Outcome: The corpus covers genres, time periods, and texts for a variety of intended audiences.
Challenges of language technologies for the indigenous languages of the Americas (C18-1)

Copied to clipboard

Challenge: Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective.
Approach: They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages .
Outcome: The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages .
The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software (2020.coling-main)

Copied to clipboard

Challenge: This paper describes the first, three-year phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in preserving their languages and extending their use.
Approach: They describe the first phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in preserving their languages.
Outcome: The proposed software will help Indigenous communities preserve and revitalize their languages and extend their use.
Modeling Northern Haida Verb Morphology (L18-1)

Copied to clipboard

Challenge: a computational model of the verbal morphology of Northern Haida is being developed . the model is capable of handling complex affixation patterns and morphophonological alternations .
Approach: They propose a computational model of the verbal morphology of Northern Haida based on finite state machines with a focus on verbs.
Outcome: The proposed model can handle complex affixation patterns and morphophonological alternations in the native language.
The Ethical Question – Use of Indigenous Corpora for Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Creating language technology based on language data is becoming more popular . indigenous language resources are not comparable in that they would encode the most recent normativised language .
Approach: They describe an ethical way to work with indigenous languages based on language data . they say data driven methods make assumptions based upon majority languages they work with . authors say data-driven methods are not ethical or beneficial .
Outcome: The proposed method is ethical and sustainable, and can be applied to indigenous languages in an ethical way.
Decolonising Speech and Language Technology (2020.coling-main)

Copied to clipboard

Challenge: Indigenous peoples are increasingly unable to go on without speech and language technologies, says a researcher . a postcolonial approach to computational methods for supporting language vitality is needed, says the researcher - lil'watul Lorna Williams .
Approach: They propose to examine colonising discourses in speech and language technology and propose a postcolonial approach to computational methods for supporting language vitality.
Outcome: The paper reviews colonising discourses in speech and language technology and suggests new ways of working with Indigenous communities.
On the Computational Modelling of Michif Verbal Morphology (2021.eacl-main)

Copied to clipboard

Challenge: Existing computational models of the verbal morphology of the Métis language are insufficient to model the language's unique phonological interactions.
Approach: They propose a finite-state computational model of the verbal morphology of Michif . they use composed finite state transducers to model concatenative morphologies .
Outcome: The proposed model is based on a series of finite-state transducers.
NLP Progress in Indigenous Latin American Languages (2024.naacl-long)

Copied to clipboard

Challenge: a new study examines the marginalization of indigenous languages in the face of rapid technological advancements.
Approach: They highlight the cultural richness of indigenous languages and the risk they face of being overlooked in the realm of natural language processing.
Outcome: The authors highlight the cultural richness of indigenous languages and their risk of being overlooked in the realm of natural language processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations